[php-src] Issue #7871: `mb_detect_encoding()` detects UTF-8 emoji byte sequence as ISO-8859-1 since PHP 8.1

From: Date: Tue, 04 Jan 2022 12:19:30 +0000
Subject: [php-src] Issue #7871: `mb_detect_encoding()` detects UTF-8 emoji byte sequence as ISO-8859-1 since PHP 8.1
Groups: php.bugs 
Request: Send a blank email to php-bugs+get-238757@lists.php.net to get a copy of this message
Issue: https://github.com/php/php-src/issues/7871 Comment Author: filecage Thanks a lot @alexdowad for looking into this! >Legacy versions of PHP return UTF-8, not because they are smart enough to tell that the string >is actually UTF-8 text, but because you put UTF-8 first in the list of candidate encodings. If you >put ISO-8859-1 first, then all versions return ISO-8859-1. When I saw that the output was different, I did not think that previous versions might be accidentally returning the correct results. >Even so, let us think a bit about this issue of emoji. It is a fact that some emoji are used by >a lot of people, and they do occupy ranges of Unicode codepoints over U+FFFF. I haven't yet had a chance to look deeply into it but I could imagine that if we're somehow able to reliably detect emojis in a string, it's a very strong indicator for a string to be UTF-8, even for short strings. It might not be worth the effort though.

« previous php.bugs (#238757) next »