Edit report at https://bugs.php.net/bug.php?id=81390&edit=1
ID: 81390
Updated by: nikic@php.net
Reported by: alec at alec dot pl
Summary: mb_detect_encoding() regression
Status: Verified
Type: Bug
Package: mbstring related
PHP Version: 8.1.0beta3
-Assigned To:
+Assigned To: alexdowad
Block user comment: N
Private report: N
New Comment:
There are multiple issues here:
1. We should consider illegal trailing characters in non-strict mode if there are still multiple
eligible encodings. Fixed in https://github.com/php/php-src/commit/43cb2548f7fd09ac3471bd71c5d28fbeaa312f2d.
This prevents detection of UCS-2 for "test:test", which has an incomplete last character.
2. Some of the "special" encodings currently don't do strict validation. uuencode is
one of those and thus ends up accepting everything. The filter should probably be fixed, though I
believe we want to drop support for these "encodings" anyway.
3. However, the real issue here is user error. You're throwing a big bucket of ambiguous
encodings at mbstring, and asking it to pick something. This kinda worked out before because in PHP
8.0 only a limited set of encodings supported encoding detection. In PHP 8.1 all encodings support
detection, including multi-byte encodings. The string "test:tes" looks nice as UTF-8, but
is also "ç¥ç´ã©´æ³" in UCS-2. Unless we want to bias detection towards
ASCII rather than CJK, both of these are sensible choices. UCS-2 is earlier in the encoding list,
and doesn't have punctuation besides.
If you want to limit detection to only UTF-8 and ISO-8859 style encodings, then you should specify
that in your encoding list. If you don't want to get back UCS-2 for something that is valid
UCS-2, don't specify UCS-2.
I think the only thing we could do here is to exclude various special encodings from detection (like
https://gist.github.com/nikic/7cab20f7286c2b9437276c4fa43f6fb4),
but apart from that I think that things are working correctly here.
Maybe Alex has some more thoughts on this.
Previous Comments:
------------------------------------------------------------------------
[2021-08-27 09:12:12] cmb@php.net
Even mb_check_encoding() fails in the same way:
<https://3v4l.org/klWd0/rfc>.
------------------------------------------------------------------------
[2021-08-26 17:20:40] alec at alec dot pl
$test = 'test:test';
$encodings = array_diff(mb_list_encodings(), ['UUENCODE', 'wchar']);
echo mb_detect_encoding($test, $encodings, true);
returns HTML-ENTITIES, this makes no sense.
------------------------------------------------------------------------
[2021-08-26 17:10:34] alec at alec dot pl
Using $test = 'test:test' in the test script also will return UUENCODE. This is really
wrong.
------------------------------------------------------------------------
[2021-08-26 17:03:35] alec at alec dot pl
Description:
------------
The same code returns "ISO-8859-1" on PHP8.0 and "UUENCODE" on PHP8.1.0beta3.
Note: the text contains some ascii with two 0x0EB characters
Note: ISO-8859-1 is before UUENCODE in the mb_list_encodings() result.
Note: Even removing UUENCODE from the list does not make it to return expected ISO-8859-1
Test script:
---------------
$test = base64_decode('Q0hBUlNFVD13aW5kb3dzLTEyNTI6RG/rO0pvaG4=');
echo mb_detect_encoding($test, mb_list_encodings());
Expected result:
----------------
ISO-8859-1
Actual result:
--------------
UUENCODE
------------------------------------------------------------------------
--
Edit this bug report at https://bugs.php.net/bug.php?id=81390&edit=1