[php-src] Issue #7871: `mb_detect_encoding()` detects UTF-8 emoji byte sequence as ISO-8859-1 since PHP 8.1
| From: | filecage | Date: | Tue, 04 Jan 2022 12:19:30 +0000 |
| Subject: | [php-src] Issue #7871: `mb_detect_encoding()` detects UTF-8 emoji byte sequence as ISO-8859-1 since PHP 8.1 | ||
| Groups: | php.bugs | ||
| Request: | Send a blank email to php-bugs+get-238757@lists.php.net to get a copy of this message | ||
Issue: https://github.com/php/php-src/issues/7871
Comment Author: filecage
Thanks a lot @alexdowad for looking into this!
>Legacy versions of PHP return UTF-8, not because they are smart enough to tell that the string
>is actually UTF-8 text, but because you put UTF-8 first in the list of candidate encodings. If you
>put ISO-8859-1 first, then all versions return ISO-8859-1.
When I saw that the output was different, I did not think that previous versions might be
accidentally returning the correct results.
>Even so, let us think a bit about this issue of emoji. It is a fact that some emoji are used by
>a lot of people, and they do occupy ranges of Unicode codepoints over U+FFFF.
I haven't yet had a chance to look deeply into it but I could imagine that if we're
somehow able to reliably detect emojis in a string, it's a very strong indicator for a string
to be UTF-8, even for short strings. It might not be worth the effort though.