note 69955 added to function.utf8-encode
| From: | php-general at lists dot php dot net | Date: | Wed, 27 Sep 2006 20:30:16 +0000 |
| Subject: | note 69955 added to function.utf8-encode | ||
| Groups: | php.notes | ||
| Request: | Send a blank email to php-notes+get-117751@lists.php.net to get a copy of this message | ||
In reply to Cundle:
Note: The BOM is completely unnecessary in UTF-8. UTF-8 is interpreted the same way regardless of
endianness, e.g. Î (U+039B, GREEK CAPITAL LETTER LAMDA) is represented as the octets 0xCE, 0x9B,
always in that order.
Extra note: UTF-16 and UCS-2 are different. The same letter would be encoded as 0x03 0x9B on
big-endian (e.g. Motorola) architecture, but 0x9B 0x03 on little-endian (e.g Intel) architecture.
But in any case, there's nothing wrong with putting a BOM at the beginning of a UTF-8 encoded
file. It is just treated as U+FEFF Zero Width No-Break Space.
----
Server IP: 64.71.164.2
Probable Submitter: 146.6.190.178
----
Manual Page -- http://www.php.net/manual/en/function.utf8-encode.php
Edit -- https://master.php.net/note/edit/69955
Del: integrated -- https://master.php.net/note/delete/69955/integrated
Del: useless -- https://master.php.net/note/delete/69955/useless
Del: bad code -- https://master.php.net/note/delete/69955/bad+code
Del: spam -- https://master.php.net/note/delete/69955/spam
Del: non-english -- https://master.php.net/note/delete/69955/non-english
Del: in docs -- https://master.php.net/note/delete/69955/in+docs
Del: other reasons-- https://master.php.net/note/delete/69955
Reject -- https://master.php.net/note/reject/69955
Search -- https://master.php.net/manage/user-notes.php