Req #31649 [Opn]: urldecode should support %uHHHH Unicode codepoint notation, which is standard

From: Date: Mon, 26 Mar 2018 21:59:18 +0000
Subject: Req #31649 [Opn]: urldecode should support %uHHHH Unicode codepoint notation, which is standard
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-214494@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=31649&edit=1 ID: 31649 Updated by: cmb@php.net Reported by: james at gogo dot co dot nz Summary: urldecode should support %uHHHH Unicode codepoint notation, which is standard Status: Open Type: Feature/Change Request Package: URL related Operating System: All PHP Version: * Block user comment: N Private report: N New Comment: The %uxxxx encoding is non-standard, and the escape() function is contained in an annex of ECMA-262 (Edition 6.0) only, which states[1]: | All of the language features and behaviours specified in this | annex have one or more undesirable characteristics and in the | absence of legacy usage would be removed from this specification. In my opinion, it does not make sense to support %uxxxx encoding in urldecode(). [1] <https://www.ecma-international.org/ecma-262/6.0/#sec-additional-ecmascript-features-for-web-browsers> Previous Comments: ------------------------------------------------------------------------ [2015-01-08 23:26:00] ajf@php.net This should probably decode to UTF-8, if it decodes to anything. ------------------------------------------------------------------------ [2005-01-21 22:46:35] derick@php.net PHP doesnt support unicode in a whole lot of places. Marking this as a feature request instead. ------------------------------------------------------------------------ [2005-01-21 22:29:07] james at gogo dot co dot nz Description: ------------ urldecode() does not understand the %uxxxx format for escaping unicode characters above 0xFF. This is a very old bug, originally reported as bug #15027 and declared bogus, I believe erroneously, and here is the reasoning... In all modern browsers (including Mozilla), JavaScript's escape() function uses %HH for Unicode codepoints below 0x0100, but %uHHHH for codepoints above there. From ECMA-262: -------------- For characters whose Unicode encoding is 0xFF or less, a two-digit escape sequence of the form %xx is used in accordance with RFC1738. For characters whose Unicode encoding is greater than 0xFF, a four-digit escape sequence of the form %uxxxx is used. -------------- I believe this is a bug, PHP is unable to urldecode the valid escape()d values from modern browsers when those escape()d strings contain unicode characters greater than 0xFF. Declaring it not a bug because it is not in the RFCs, but rather defined by ECMA is a poor decision. Reproduce code: --------------- echo urldecode('%u2013'); Expected result: ---------------- A string containing the three characters comprising the unicode character 0x2013 (En Dash) in utf-8, namely 0xE2 0x80 and 0x93. Actual result: -------------- The literal string "%u2013". ------------------------------------------------------------------------ -- Edit this bug report at https://bugs.php.net/bug.php?id=31649&edit=1

« previous php.bugs (#214494) next »