Edit report at https://bugs.php.net/bug.php?id=62010&edit=1
ID: 62010
Comment by: tklingenberg at lastflood dot net
Reported by: tklingenberg at lastflood dot net
Summary: json_decode produces invalid byte-sequences
Status: Assigned
Type: Bug
Package: JSON related
Operating System: Windows
PHP Version: 5.3.13
Assigned To: bukka
Block user comment: N
Private report: N
New Comment:
Hi bukka,
thank you for taking the time to look into it.
But I'm very sorry to highlight that the information you've provided in your comment is
best of all only remotely related to this issue and does not touch the root-cause of the flaw
reported here.
You've perhaps been misguided by the internals mailings (haven't read those), the part you
quote is about binary string data.
But the report I created is *not* about binary data, you can see, the string presented is US-ASCII
without any control characters.
It's about the strings _represented_ by JSON (not in binary), and more specifically the option
to use an \uXXXX (six characters) escape sequence for any character in the Basic Multilingual Plane
(U+0000 through U+FFFF). Please see Section 7 of the JSON RFC.
U+D834 is not a character in the Basic Multilingual Plane (see Unicode, compare with a reference,
exemplary: http://www.fileformat.info/info/unicode/char/d834/index.htm).
If a string would have been passed json_decode containing the related binary sequence - as what you
say would be allowed by the JSON spec - PHP handles it correctly according the documented contract:
The binary sequence would qualify as *not* being an UTF-8 string and therefore the result of the
function is unexpected:
<?php
$notUTF8 = "\"\xED\xA0\xB4\"";
var_dump($notUTF8); // string(5) ""���""
$result = json_decode($notUTF8);
var_dump($result); // NULL
That's covered by the specs you quote, but not the flaw I reported here. As you can see, this
is a different example and I can't see that PHP violates the spec nor it's own contract
here.
Previous Comments:
------------------------------------------------------------------------
[2015-05-28 18:58:45] bukka@php.net
I just emailed about this on internals. This is not a bug as it is conformant with the JSON RFC 7159
as noted in section 8.2:
However, the ABNF in this specification allows member names and
string values to contain bit sequences that cannot encode Unicode
characters; for example, "\uDEAD" (a single unpaired UTF-16
surrogate). Instances of this have been observed, for example, when
a library truncates a UTF-16 string without checking whether the
truncation split a surrogate pair. The behavior of software that
receives JSON texts containing such values is unpredictable; for
example, implementations might return different values for the length
of a string value or even suffer fatal runtime exceptions.
As you can see that behavior is unpredictable.
However I see a use case here and that's why I proposed new option JSON_VALID_ESCAPED_UNICODE
that would emit JSON_ERROR_UTF16 if such sequence appears in decoded string.
------------------------------------------------------------------------
[2013-07-12 15:55:24] masakielastic at gmail dot com
Here is RFC 3629's description about UTF-8 definition.
The definition of UTF-8 prohibits encoding character numbers
between U+D800 and U+DFFF, which are reserved for use with the
UTF-16 encoding form (as surrogate pairs) and do not directly
represent characters.
http://tools.ietf.org/html/rfc3629
The following patch solve the part of problem,
The isolated low surrogate pairs(U+DC00 U+DFFF) are replaced with U+FFFD,
The imrovement for high surrogate pairs (U+D800 - U+DBFF) is needed.
https://gist.github.com/masakielastic/5985383
var_dump(
"\xef\xbf\xbd" === json_decode('"\udc00"'),
"\xef\xbf\xbd"."\xed\xa0\x80" ===
json_decode('"\ud800\ud800"'),
"\xed\xa0\x80" === json_decode('"\ud800"')
);
The consistency for the following options
(under the discussion) is needed too.
json_encode's option for replacing ill-formd byte sequences
with substitute characters
https://bugs.php.net/bug.php?id=65082
------------------------------------------------------------------------
[2013-01-11 09:44:55] votefordevnull at gmail dot com
Successfully reproduced on Linux
------------------------------------------------------------------------
[2012-05-11 22:46:34] tklingenberg at lastflood dot net
Looks like that #41067 https://bugs.php.net/bug.php?id=41067 was not fully
fixed.
------------------------------------------------------------------------
[2012-05-11 22:12:42] tklingenberg at lastflood dot net
Description:
------------
It's a typical case the JSON *and* UTF-16 specifications warn about: decoding of
non-existing UTF-16 code-points:
json_decode('"\ud834"')
shoud give NULL because \ud834 is *invalid*. But instead it starts some party,
get's boozed and offers this as UTF-8 byte-sequence:
1110 1101 1010 0000 1011 0100
1110 xxxx 10xx xxxx 10xx xxxx
1101 1000 0011 0100
D8 34
U+D834 is not a valid unicode character.
Test script:
---------------
if (NULL !== json_decode('"\ud834"')) {
echo "json_decode is still broken.";
}
Expected result:
----------------
NULL because the json is invalid.
Actual result:
--------------
PHP tries to create UTF-8 out of it and fails by creating invalid UTF-8 unicode
byte-sequences.
------------------------------------------------------------------------
--
Edit this bug report at https://bugs.php.net/bug.php?id=62010&edit=1