#43896 [Opn->Fbk]: htmlspecialchars returns empty string on invalid unicode sequence
| From: | jani@php.net | Date: | Sun, 27 Jul 2008 20:24:10 +0000 |
| Subject: | #43896 [Opn->Fbk]: htmlspecialchars returns empty string on invalid unicode sequence | ||
| References: | 1 | Groups: | php.bugs |
| Request: | Send a blank email to php-bugs+get-127329@lists.php.net to get a copy of this message | ||
ID: 43896
Updated by: jani@php.net
Reported By: arnaud dot lb at gmail dot com
-Status: Open
+Status: Feedback
Bug Type: Strings related
Operating System: *
PHP Version: 5.2CVS, 5.3CVS (2008-07-15)
Previous Comments:
------------------------------------------------------------------------
[2008-07-18 00:10:45] moriyoshi@php.net
I even don't think this is a valid bug in the first place. You passed a
string that is encoded in ISO-8859-15 to htmlspecialchars() while
specifying UTF-8 to force the string to be treated as "UTF-8". One
should never depend on the past wrond behaviour with which invalid byte
sequences pass through. Besides, you can always work around it by
giving
ISO-8859-15 to the third argument.
------------------------------------------------------------------------
[2008-06-27 17:32:43] sillyxone at yaoo dot com
is also affected in 5.2, for example:
$str = 'Hello' . chr(160) . 'there';
print(htmlentities($str, ENT_COMPAT, 'UTF-8'));
Instead of printing "Hello there", it prints nothing (empty string).
The same for htmlspecialchars().
Both functions work fine in 5.1
------------------------------------------------------------------------
[2008-05-05 21:00:37] heurika at gmail dot com
Hi,
I've got the same Bug, posted on #43740.
Please fix it.
Thanks!
------------------------------------------------------------------------
[2008-02-17 13:25:22] andreas dot ravnestad at gmail dot com
This seems to be breaking PEAR::Text_Wiki completely when using UTF-8:
http://pear.php.net/bugs/bug.php?id=13136
------------------------------------------------------------------------
[2008-01-24 20:51:11] tallyce at gmail dot com
See also bugs 43294 and 43549 which seem to be the same thing.
This is really starting to bite now. Please can this be fixed, or
suggest how we can reliably process incoming user data in UTF8 given
this behaviour change!
------------------------------------------------------------------------
The remainder of the comments for this report are too long. To view
the rest of the comments, please view the bug report online at
http://bugs.php.net/43896
--
Edit this bug report at http://bugs.php.net/?id=43896&edit=1