note 69955 deleted from function.utf8-encode by cmb
| From: | cmb@php.net | Date: | Sun, 22 Dec 2019 08:58:57 +0000 |
| Subject: | note 69955 deleted from function.utf8-encode by cmb | ||
| References: | 1 | Groups: | php.notes |
| Request: | Send a blank email to php-notes+get-213572@lists.php.net to get a copy of this message | ||
Note Submitter:
----
In reply to Cundle:
Note: The BOM is completely unnecessary in UTF-8. UTF-8 is interpreted the same way regardless of
endianness, e.g. Î (U+039B, GREEK CAPITAL LETTER LAMDA) is represented as the octets 0xCE, 0x9B,
always in that order.
Extra note: UTF-16 and UCS-2 are different. The same letter would be encoded as 0x03 0x9B on
big-endian (e.g. Motorola) architecture, but 0x9B 0x03 on little-endian (e.g Intel) architecture.
But in any case, there's nothing wrong with putting a BOM at the beginning of a UTF-8 encoded
file. It is just treated as U+FEFF Zero Width No-Break Space.