Edit report at https://bugs.php.net/bug.php?id=62010&edit=1
ID: 62010
Updated by: bukka@php.net
Reported by: tklingenberg at lastflood dot net
Summary: json_decode produces invalid byte-sequences
-Status: Open
+Status: Assigned
Type: Bug
Package: JSON related
Operating System: Windows
PHP Version: 5.3.13
-Assigned To:
+Assigned To: bukka
Block user comment: N
Private report: N
New Comment:
I just emailed about this on internals. This is not a bug as it is conformant with the JSON RFC 7159
as noted in section 8.2:
However, the ABNF in this specification allows member names and
string values to contain bit sequences that cannot encode Unicode
characters; for example, "\uDEAD" (a single unpaired UTF-16
surrogate). Instances of this have been observed, for example, when
a library truncates a UTF-16 string without checking whether the
truncation split a surrogate pair. The behavior of software that
receives JSON texts containing such values is unpredictable; for
example, implementations might return different values for the length
of a string value or even suffer fatal runtime exceptions.
As you can see that behavior is unpredictable.
However I see a use case here and that's why I proposed new option JSON_VALID_ESCAPED_UNICODE
that would emit JSON_ERROR_UTF16 if such sequence appears in decoded string.
Previous Comments:
------------------------------------------------------------------------
[2013-07-12 15:55:24] masakielastic at gmail dot com
Here is RFC 3629's description about UTF-8 definition.
The definition of UTF-8 prohibits encoding character numbers
between U+D800 and U+DFFF, which are reserved for use with the
UTF-16 encoding form (as surrogate pairs) and do not directly
represent characters.
http://tools.ietf.org/html/rfc3629
The following patch solve the part of problem,
The isolated low surrogate pairs(U+DC00 U+DFFF) are replaced with U+FFFD,
The imrovement for high surrogate pairs (U+D800 - U+DBFF) is needed.
https://gist.github.com/masakielastic/5985383
var_dump(
"\xef\xbf\xbd" === json_decode('"\udc00"'),
"\xef\xbf\xbd"."\xed\xa0\x80" ===
json_decode('"\ud800\ud800"'),
"\xed\xa0\x80" === json_decode('"\ud800"')
);
The consistency for the following options
(under the discussion) is needed too.
json_encode's option for replacing ill-formd byte sequences
with substitute characters
https://bugs.php.net/bug.php?id=65082
------------------------------------------------------------------------
[2013-01-11 09:44:55] votefordevnull at gmail dot com
Successfully reproduced on Linux
------------------------------------------------------------------------
[2012-05-11 22:46:34] tklingenberg at lastflood dot net
Looks like that #41067 https://bugs.php.net/bug.php?id=41067 was not fully
fixed.
------------------------------------------------------------------------
[2012-05-11 22:12:42] tklingenberg at lastflood dot net
Description:
------------
It's a typical case the JSON *and* UTF-16 specifications warn about: decoding of
non-existing UTF-16 code-points:
json_decode('"\ud834"')
shoud give NULL because \ud834 is *invalid*. But instead it starts some party,
get's boozed and offers this as UTF-8 byte-sequence:
1110 1101 1010 0000 1011 0100
1110 xxxx 10xx xxxx 10xx xxxx
1101 1000 0011 0100
D8 34
U+D834 is not a valid unicode character.
Test script:
---------------
if (NULL !== json_decode('"\ud834"')) {
echo "json_decode is still broken.";
}
Expected result:
----------------
NULL because the json is invalid.
Actual result:
--------------
PHP tries to create UTF-8 out of it and fails by creating invalid UTF-8 unicode
byte-sequences.
------------------------------------------------------------------------
--
Edit this bug report at https://bugs.php.net/bug.php?id=62010&edit=1