Bug #81390 [Asn]: mb_detect_encoding() regression

From: Date: Sun, 05 Dec 2021 16:01:08 +0000
Subject: Bug #81390 [Asn]: mb_detect_encoding() regression
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-238194@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=81390&edit=1

 ID:                 81390
 Updated by:         alexdowad@php.net
 Reported by:        alec at alec dot pl
 Summary:            mb_detect_encoding() regression
 Status:             Assigned
 Type:               Bug
 Package:            mbstring related
 PHP Version:        8.1.0beta5
 Assigned To:        alexdowad
 Block user comment: N
 Private report:     N

 New Comment:

Hi, Alec. Please keep up the great work with testing. I tried the text which you linked to:

---------------
$text = "Ola szykuje się do szkoły. Jest już w piątej klasie. Dawniej obawiała
się szkoły, teraz bardzo lubi tam chodzić. W szkole nie tylko uczy się ciekawych rzeczy
– spotyka też swoich kolegów i koleżanki. Najbardziej lubi spędzać czas ze
swoimi przyjaciółmi z klasy – są to Monika i Michał.

Ola lubi wszystkie przedmioty. Wie, że nauka jest ważna. Najmilej spędza czas na lekcjach o
przyrodzie – Ola bardzo lubi zwierzęta. W klasie Oli mieszka chomik. Wszystkie dzieci
dbają o niego. Przynoszą mu jedzenie i głaszczą. Ola nie ma własnego zwierzęcia,
więc chomik to kolejny powód dla którego lubi chodzić do szkoły.

Szkoła Oli jest blisko jej domu. Ola chodzi do szkoły sama. Gdy Ola była młodsza –
zawsze odprowadzał ją tam tata, mama lub babcia. Przed każdym dniem szkoły Ola pakuje
plecak. Sprawdza czy są wszystkie książki i zeszyty. Sprawdza też piórnik – czy
jest tam linijka, długopis, ołówek, gumka i mazaki.

W przerwie między lekcjami Ola je w szkole obiad. Dzięki temu ma dużo siły na dalszą
naukę. Po szkole droga do domu jest krótka, ale zajmuje Oli dużo czasu. Dlaczego? Bo to
czas, w którym lubi rozmawiać z koleżankami i kolegami podczas spaceru do domu. Ola jednak
wie, że nie może wrócić późno.";

echo mb_detect_encoding($text,'utf-8,windows-1252,iso-8859-1,iso-8859-2', true),
"\n";
-------------

And then, in my shell:

-----------------
17:56 ~prog/php/php-src % ./sapi/cli/php test.php
UTF-8
-----------------

Looks good to me! That's on the tip of my development branch, which (crucially) does include
the latest patch for Polish... the one which you said you didn't use.

Please do your best to see if you can still find breakage anywhere! If you find anything else which
is arguably a bug, I will be very, very happy to fix it!


Previous Comments:
------------------------------------------------------------------------
[2021-11-30 12:28:24] cmb@php.net

Related To: Bug #81676

------------------------------------------------------------------------
[2021-11-28 10:38:30] alec at alec dot pl

I've another case which is either a bug or an indication that the method can't really
distinguish utf-8 from iso-8859-*/windows-125X charsets. I'm not sure it's only a Polish
thing, but most likely not.

I took a sample text from https://lingua.com/pl/polski/czytanie/czas-do-szkoly/
and used it with

mb_detect_encoding($text,'utf-8,windows-1252,iso-8859-1,iso-8859-2', true);

It returns windows-1252, but it should be utf-8. If I convert the text to iso-8859-2, it still
returns windows-1252.

ps. I didn't use your latest patch for Polish.

------------------------------------------------------------------------
[2021-11-25 08:42:10] alexdowad@php.net

Thanks, Alec.

Hopefully this will be merged soon:

https://github.com/php/php-src/pull/7659/commits/889d2fe3bafdb8e9da39b2297d6e9f13a518584e

------------------------------------------------------------------------
[2021-11-25 06:32:31] alec at alec dot pl

Well, let's take Polish language (which is ISO-8859-2 family) as an example. The one I know.

There's a table https://pl.wikipedia.org/wiki/Alfabet_polski#Cz%C4%99sto%C5%9B%C4%87_wyst%C4%99powania_liter
(only in Polish version). According to this, some diacritical characters are quite common (~1-2%).
E.g. letter Ł is more common that letter B.

I have no idea how that applies to the whole character set detection, though.

------------------------------------------------------------------------
[2021-11-24 20:05:19] alexdowad@php.net

In answer to Alec:

> I'm guessing that in earlier versions there was no such thing as demerits (or the algo was
> even more different), so the provided priority list had more impact on the result. Is that right?

Yep. The earlier versions of mb_detect_encoding only checked which encodings 
the input string was valid in, and picked the first one on the list. If the string was valid in more
than one encoding, it did not do anything at all to try to figure out which one was most likely.

It also was not able to detect every text encoding supported by mbstring, only some of them.

> I'm not that good with the subject, but maybe you can get some ideas from https://github.com/Joungkyun/libchardet or https://github.com/CLD2Owners/cld2

Thanks for the references! I briefly browsed a bit of the code for libchardet. It looks like they
also rely heavily on character frequency tables... but instead of having just one table of
'common' and 'rare' characters, like mbstring, they have tables *for each
supported text encoding*. So they are able to say that "this codepoint rarely appears in EUC-JP
encoded text", or "this codepoint often occurs in UTF-16 encoded text".

It looks like they have some other tricks as well, though I didn't read the code in enough
detail to figure them all out.

One other thing here. It looks like there is already a PHP extension for libchardet. Therefore, I
don't think there is much need to compete with them; if people want more sophisticated charset
detection in PHP, they can just use libchardet.

Still, if there is anything we can do with a reasonable amount of effort to improve detection
accuracy across the board, I am open to ideas. If someone is willing to pitch in and help with some
of the work, that would be great.

You are also very right that misdetection is much more likely on short strings than on long ones.

> ps. the input is proper iso-8859-2 string, so why such a big difference between these?:

You can see the ranges of codepoints which are currently considered 'common' by mbstring
here:

https://github.com/php/php-src/blob/master/ext/mbstring/common_codepoints.txt

You will notice that U+0144 (ń) is not there. Actually, I haven't included anything from the
"Latin Extended-A" range. You can see all the characters in this range here:

https://www.unicode.org/charts/PDF/U0100.pdf

If you think any of these should be considered 'common', please let us know and we can
tweak the table accordingly.

------------------------------------------------------------------------


The remainder of the comments for this report are too long. To view
the rest of the comments, please view the bug report online at

    https://bugs.php.net/bug.php?id=81390


--
Edit this bug report at https://bugs.php.net/bug.php?id=81390&edit=1


Thread (37 messages)

« previous php.bugs (#238194) next »