#23449 [Opn]: php doesn't ignore the utf-8 BOM
| From: | brofield at jellycan dot com | Date: | Fri, 02 May 2003 10:28:54 +0000 |
| Subject: | #23449 [Opn]: php doesn't ignore the utf-8 BOM | ||
| References: | 1 | Groups: | php.bugs |
| Request: | Send a blank email to php-bugs+get-38937@lists.php.net to get a copy of this message | ||
ID: 23449
User updated by: brofield at jellycan dot com
-Summary: htmlentities uses wrong entity for U+2225
Reported By: brofield at jellycan dot com
Status: Open
Bug Type: Strings related
Operating System: windows 2000
PHP Version: 4.3.1
New Comment:
Calling syntax is...
$str = htmlentities( "string containing U+2225", ENT_QUOTES, "utf-8"
);
Other incorrect conversions seem to abound...
$str = ctwEncodeUtf8(
"∥∦∧∨∩∪∫" );
echo $str.'<br>';
$str = htmlentities( $str, ENT_QUOTES, "utf-8" );
echo $str.'<br>';
echo htmlentities( $str, ENT_QUOTES, 'utf-8' ).'<br>';
exit;
Note: ctwEncodeUtf8 is the same function as found at
http://www.zend.com/codex.php?id=838&single=1
Results:
a∦ÈÉ¿¾ç
¿¾çÉ¿¾ç
∩∪∫É¿¾ç
The first and second lines should be the same.
Previous Comments:
------------------------------------------------------------------------
[2003-05-02 05:21:18] brofield at jellycan dot com
There seems to be a bug in htmlentities() utf-8 conversion. The unicode
character U+2225 a gets converted into ∩ which is the html entity
for the unicode character U+2229 ¿. It should be using ∥ to get
the correct symbol.
------------------------------------------------------------------------
--
Edit this bug report at http://bugs.php.net/?id=23449&edit=1