[php-src] Issue #8279: mb_detect_encoding does not return the first matching encoding anymore

From: Date: Mon, 25 Apr 2022 14:27:08 +0000
Subject: [php-src] Issue #8279: mb_detect_encoding does not return the first matching encoding anymore
Groups: php.bugs 
Request: Send a blank email to php-bugs+get-241228@lists.php.net to get a copy of this message
Issue: https://github.com/php/php-src/issues/8279 Comment Author: come-nc > I don't know if that accented letter is commonly used in any major language of the world, > but currently mb_detect_encoding classifies it as a "rare" > character and in this case, penalizes UTF-8 because of it. If you remove ŷ from that string, > mb_detect_encoding will decide that UTF-8 is more likely what you > wanted. > > If the input string was longer, then UTF-8 would have much better chances of winning out. In > general, trying to automatically detect text encoding on short strings is very error-prone. This is > also true for the simpler approach of "just picking an encoding that works"; as the input > string becomes shorter, the chances that an unintended encoding will match by accident become > higher. > > If there are good reasons why mb_detect_encoding should > consider ŷ to be a "common" character, please feel free to propose that. It was reported to us that the problem was happening on common slavic names. I tried several names from https://en.wikipedia.org/wiki/Slavic_names and found that both 'Dušan' and 'Živko' are wrongly detected as latin1. I understand that detecting encoding on a short string is hard but it is a problem that the function is failing on common words, and it should in this case use the order of the passed encoding to return the first valid one as before. See https://3v4l.org/dX7FR

« previous php.bugs (#241228) next »