Edit report at https://bugs.php.net/bug.php?id=81390&edit=1
ID: 81390
Updated by: alexdowad@php.net
Reported by: alec at alec dot pl
Summary: mb_detect_encoding() regression
Status: Assigned
Type: Bug
Package: mbstring related
PHP Version: 8.1.0beta5
Assigned To: alexdowad
Block user comment: N
Private report: N
New Comment:
Hi, Alec. Please keep up the great work with testing. I tried the text which you linked to:
---------------
$text = "Ola szykuje siÄ do szkoÅy. Jest już w piÄ
tej klasie. Dawniej obawiaÅa
siÄ szkoÅy, teraz bardzo lubi tam chodziÄ. W szkole nie tylko uczy siÄ ciekawych rzeczy
â spotyka też swoich kolegów i koleżanki. Najbardziej lubi spÄdzaÄ czas ze
swoimi przyjacióÅmi z klasy â sÄ
to Monika i MichaÅ.
Ola lubi wszystkie przedmioty. Wie, że nauka jest ważna. Najmilej spÄdza czas na lekcjach o
przyrodzie â Ola bardzo lubi zwierzÄta. W klasie Oli mieszka chomik. Wszystkie dzieci
dbajÄ
o niego. PrzynoszÄ
mu jedzenie i gÅaszczÄ
. Ola nie ma wÅasnego zwierzÄcia,
wiÄc chomik to kolejny powód dla którego lubi chodziÄ do szkoÅy.
SzkoÅa Oli jest blisko jej domu. Ola chodzi do szkoÅy sama. Gdy Ola byÅa mÅodsza â
zawsze odprowadzaÅ jÄ
tam tata, mama lub babcia. Przed każdym dniem szkoÅy Ola pakuje
plecak. Sprawdza czy sÄ
wszystkie ksiÄ
żki i zeszyty. Sprawdza też piórnik â czy
jest tam linijka, dÅugopis, oÅówek, gumka i mazaki.
W przerwie miÄdzy lekcjami Ola je w szkole obiad. DziÄki temu ma dużo siÅy na dalszÄ
naukÄ. Po szkole droga do domu jest krótka, ale zajmuje Oli dużo czasu. Dlaczego? Bo to
czas, w którym lubi rozmawiaÄ z koleżankami i kolegami podczas spaceru do domu. Ola jednak
wie, że nie może wróciÄ późno.";
echo mb_detect_encoding($text,'utf-8,windows-1252,iso-8859-1,iso-8859-2', true),
"\n";
-------------
And then, in my shell:
-----------------
17:56 ~prog/php/php-src % ./sapi/cli/php test.php
UTF-8
-----------------
Looks good to me! That's on the tip of my development branch, which (crucially) does include
the latest patch for Polish... the one which you said you didn't use.
Please do your best to see if you can still find breakage anywhere! If you find anything else which
is arguably a bug, I will be very, very happy to fix it!
Previous Comments:
------------------------------------------------------------------------
[2021-11-30 12:28:24] cmb@php.net
Related To: Bug #81676
------------------------------------------------------------------------
[2021-11-28 10:38:30] alec at alec dot pl
I've another case which is either a bug or an indication that the method can't really
distinguish utf-8 from iso-8859-*/windows-125X charsets. I'm not sure it's only a Polish
thing, but most likely not.
I took a sample text from https://lingua.com/pl/polski/czytanie/czas-do-szkoly/
and used it with
mb_detect_encoding($text,'utf-8,windows-1252,iso-8859-1,iso-8859-2', true);
It returns windows-1252, but it should be utf-8. If I convert the text to iso-8859-2, it still
returns windows-1252.
ps. I didn't use your latest patch for Polish.
------------------------------------------------------------------------
[2021-11-25 08:42:10] alexdowad@php.net
Thanks, Alec.
Hopefully this will be merged soon:
https://github.com/php/php-src/pull/7659/commits/889d2fe3bafdb8e9da39b2297d6e9f13a518584e
------------------------------------------------------------------------
[2021-11-25 06:32:31] alec at alec dot pl
Well, let's take Polish language (which is ISO-8859-2 family) as an example. The one I know.
There's a table https://pl.wikipedia.org/wiki/Alfabet_polski#Cz%C4%99sto%C5%9B%C4%87_wyst%C4%99powania_liter
(only in Polish version). According to this, some diacritical characters are quite common (~1-2%).
E.g. letter Å is more common that letter B.
I have no idea how that applies to the whole character set detection, though.
------------------------------------------------------------------------
[2021-11-24 20:05:19] alexdowad@php.net
In answer to Alec:
> I'm guessing that in earlier versions there was no such thing as demerits (or the algo was
> even more different), so the provided priority list had more impact on the result. Is that right?
Yep. The earlier versions of mb_detect_encoding only checked which encodings
the input string was valid in, and picked the first one on the list. If the string was valid in more
than one encoding, it did not do anything at all to try to figure out which one was most likely.
It also was not able to detect every text encoding supported by mbstring, only some of them.
> I'm not that good with the subject, but maybe you can get some ideas from https://github.com/Joungkyun/libchardet or https://github.com/CLD2Owners/cld2
Thanks for the references! I briefly browsed a bit of the code for libchardet. It looks like they
also rely heavily on character frequency tables... but instead of having just one table of
'common' and 'rare' characters, like mbstring, they have tables *for each
supported text encoding*. So they are able to say that "this codepoint rarely appears in EUC-JP
encoded text", or "this codepoint often occurs in UTF-16 encoded text".
It looks like they have some other tricks as well, though I didn't read the code in enough
detail to figure them all out.
One other thing here. It looks like there is already a PHP extension for libchardet. Therefore, I
don't think there is much need to compete with them; if people want more sophisticated charset
detection in PHP, they can just use libchardet.
Still, if there is anything we can do with a reasonable amount of effort to improve detection
accuracy across the board, I am open to ideas. If someone is willing to pitch in and help with some
of the work, that would be great.
You are also very right that misdetection is much more likely on short strings than on long ones.
> ps. the input is proper iso-8859-2 string, so why such a big difference between these?:
You can see the ranges of codepoints which are currently considered 'common' by mbstring
here:
https://github.com/php/php-src/blob/master/ext/mbstring/common_codepoints.txt
You will notice that U+0144 (Å) is not there. Actually, I haven't included anything from the
"Latin Extended-A" range. You can see all the characters in this range here:
https://www.unicode.org/charts/PDF/U0100.pdf
If you think any of these should be considered 'common', please let us know and we can
tweak the table accordingly.
------------------------------------------------------------------------
The remainder of the comments for this report are too long. To view
the rest of the comments, please view the bug report online at
https://bugs.php.net/bug.php?id=81390
--
Edit this bug report at https://bugs.php.net/bug.php?id=81390&edit=1