[php-src] Issue #8279: mb_detect_encoding does not return the first matching encoding anymore
| From: | come-nc | Date: | Mon, 25 Apr 2022 14:27:08 +0000 |
| Subject: | [php-src] Issue #8279: mb_detect_encoding does not return the first matching encoding anymore | ||
| Groups: | php.bugs | ||
| Request: | Send a blank email to php-bugs+get-241228@lists.php.net to get a copy of this message | ||
Issue: https://github.com/php/php-src/issues/8279
Comment Author: come-nc
> I don't know if that accented letter is commonly used in any major language of the world,
> but currently
mb_detect_encoding classifies it as a "rare"
> character and in this case, penalizes UTF-8 because of it. If you remove ŷ from that string,
> mb_detect_encoding will decide that UTF-8 is more likely what you
> wanted.
>
> If the input string was longer, then UTF-8 would have much better chances of winning out. In
> general, trying to automatically detect text encoding on short strings is very error-prone. This is
> also true for the simpler approach of "just picking an encoding that works"; as the input
> string becomes shorter, the chances that an unintended encoding will match by accident become
> higher.
>
> If there are good reasons why mb_detect_encoding should
> consider ŷ to be a "common" character, please feel free to propose that.
It was reported to us that the problem was happening on common slavic names. I tried several names
from https://en.wikipedia.org/wiki/Slavic_names and
found that both 'Dušan' and 'Živko' are wrongly detected as latin1.
I understand that detecting encoding on a short string is hard but it is a problem that the function
is failing on common words, and it should in this case use the order of the passed encoding to
return the first valid one as before.
See https://3v4l.org/dX7FR