Bug #81390 [Csd->Asn]: mb_detect_encoding() regression

From: Date: Sat, 25 Sep 2021 12:36:25 +0000
Subject: Bug #81390 [Csd->Asn]: mb_detect_encoding() regression
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-236829@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=81390&edit=1 ID: 81390 User updated by: alec at alec dot pl Reported by: alec at alec dot pl Summary: mb_detect_encoding() regression -Status: Closed +Status: Assigned Type: Bug Package: mbstring related PHP Version: 8.1.0beta3 Assigned To: alexdowad Block user comment: N Private report: N New Comment: Can I have a clarification on which version this is fixed? Previous Comments: ------------------------------------------------------------------------ [2021-09-24 06:12:07] alec at alec dot pl The original case is still not fixed in 8.1.0rc2. ------------------------------------------------------------------------ [2021-09-06 19:55:45] alexdowad@php.net Will submit the fix in my next PR for mbstring. ------------------------------------------------------------------------ [2021-09-06 19:43:43] alexdowad@php.net OK, just found and fixed a bug! After the fix, Alec's test case returns ISO-8859-1. Thanks for helping us to find it, Alec! ------------------------------------------------------------------------ [2021-08-27 17:46:20] alexdowad@php.net Thanks to Alec for the report! Some comments: The new legacy encoding detection code, which (as Nikita mentioned) is intended to work with all supported encodings, uses a couple of simple heuristics: - If the input string is not valid in a candidate encoding, that encoding is immediately rejected. - When the input string is converted to a candidate encoding, each control character or codepoint in Unicode's Private Use Area counts for 10 "demerits" against the candidate - Each punctuation character counts for 1 "demerit", since punctuation is much less common in natural language strings than letters (also, when Shift-JIS or ISO-2022 strings are misinterpreted as ASCII, they tend to have large numbers of punctuation characters). We can easily add more heuristics, and refine the existing ones. We will _never_ get anything close to 100% accuracy; frankly, even with human intelligence, it is not always possible to figure out what the intended encoding of some random string is. We should favor heuristics which will improve detection accuracy in a wide range of situations, and which are fast to evaluate. Here are a couple of ideas: - Consider completely banning uuencode, base64, QPrint, 'HTML entities', '7-bit', and '8-bit' from being returned as the detected text encoding. - (Line 46 of mbfl_encoding.h is interesting; it shows that the original author of MBString recognized that these are not really 'text encodings' in the same sense that the other supported encodings are.) - Rank the other supported encodings according to the likelihood that they will be encountered 'in the wild', and favor those higher on the list when more than one candidate is possible. Here's another thought. Two common scenarios when encoding detection returns the wrong result: 1) The string is changed into all or almost all CJK characters, which are usually _very rare_ CJK characters. 2) (This is what happens when a CJK string is mistakenly detected as being ASCII or a 'european' encoding) The string is changed into a mishmash of letters and punctuation, with very few spaces. This implies that it might be helpful to classify CJK characters as 'common' and 'rare', perhaps using something like a Bloom filter. For strings with a high proportion of 'european' characters, maybe we should expect a good number of spaces (unless the string is just a single word). I think that finding PUA codepoints is a fairly good indicator that a candidate encoding might be wrong, but rather than the current test, we could simply check for a range of values (i.e. c >= PUA_MIN && c <= PUA_MAX). This might help to speed things up, since we are also seeing reports that the new encoding detection code is too slow for some users. ------------------------------------------------------------------------ [2021-08-27 17:02:28] alec at alec dot pl I guess I'll use a sane list of encodings, however there's still somethings wrong. $test = 'test:test'; $encodings = ['UTF-8', 'SJIS', 'GB2312', 'ISO-8859-1', 'ISO-8859-2', 'ISO-8859-3', 'ISO-8859-4', 'ISO-8859-5', 'ISO-8859-6', 'ISO-8859-7', 'ISO-8859-8', 'ISO-8859-9', 'ISO-8859-10', 'ISO-8859-13', 'ISO-8859-14', 'ISO-8859-15', 'ISO-8859-16', 'WINDOWS-1252', 'WINDOWS-1251', 'EUC-JP', 'EUC-TW', 'KOI8-R', 'BIG-5', 'ISO-2022-KR', 'ISO-2022-JP', 'UTF-16' ]; echo mb_detect_encoding($test, $encodings); returns "UTF-16". Maybe that's one of the issues you described already. ------------------------------------------------------------------------ The remainder of the comments for this report are too long. To view the rest of the comments, please view the bug report online at https://bugs.php.net/bug.php?id=81390 -- Edit this bug report at https://bugs.php.net/bug.php?id=81390&edit=1

« previous php.bugs (#236829) next »