Bug #81390 [Com]: mb_detect_encoding() regression

From: Date: Tue, 23 Nov 2021 07:48:36 +0000
Subject: Bug #81390 [Com]: mb_detect_encoding() regression
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-237932@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=81390&edit=1

 ID:                 81390
 Comment by:         alec at alec dot pl
 Reported by:        alec at alec dot pl
 Summary:            mb_detect_encoding() regression
 Status:             Assigned
 Type:               Bug
 Package:            mbstring related
 PHP Version:        8.1.0beta5
 Assigned To:        alexdowad
 Block user comment: N
 Private report:     N

 New Comment:

So, this is what I expected. Short input can produce unexpected results. I'm guessing that in
earlier versions there was no such thing as demerits (or the algo was even more different), so the
provided priority list had more impact on the result. Is that right?

I'm not sure we should consider this last case a bug anymore.

I'm not that good with the subject, but maybe you can get some ideas from https://github.com/Joungkyun/libchardet or https://github.com/CLD2Owners/cld2

ps. the input is proper iso-8859-2 string, so why such a big difference between these?:
Score for ISO-8859-2: 0 illegal chars, 37 demerits
Score for ISO-8859-1: 0 illegal chars, 8 demerits


Previous Comments:
------------------------------------------------------------------------
[2021-11-22 18:24:28] alexdowad@php.net

Patrick Allaert reached out to me today to see if the latest report from Alec could be checked into
before he cuts the final release for 8.1.0. Thanks very much for the 'heads up', Patrick!
Gladly!

I added the following line to mbfl_encoding_detector_judge in mbfilter.c and
recompiled:

    printf("Score for %s: %d illegal chars, %d demerits\n", filter->from->name,
data->num_illegalchars, data->score);

And then ran Alec's new test case. Output:

    Score for UTF-8: 1 illegal chars, 5 demerits                                                    
 
    Score for ISO-8859-1: 0 illegal chars, 8 demerits                                               
 
    Score for ISO-8859-2: 0 illegal chars, 37 demerits                                              
 
    Score for ISO-8859-3: 0 illegal chars, 8 demerits                                               
 
    Score for ISO-8859-4: 0 illegal chars, 37 demerits                                              
 
    Score for ISO-8859-5: 0 illegal chars, 8 demerits 
    Score for ISO-8859-6: 0 illegal chars, 8 demerits 
    Score for ISO-8859-7: 0 illegal chars, 8 demerits 
    Score for ISO-8859-8: 0 illegal chars, 8 demerits 
    Score for ISO-8859-9: 0 illegal chars, 8 demerits 
    Score for ISO-8859-10: 0 illegal chars, 37 demerits
    Score for ISO-8859-13: 0 illegal chars, 37 demerits
    Score for ISO-8859-14: 0 illegal chars, 8 demerits
    Score for ISO-8859-15: 0 illegal chars, 8 demerits
    Score for ISO-8859-16: 0 illegal chars, 37 demerits
    Score for Windows-1252: 0 illegal chars, 8 demerits    
    Score for Windows-1251: 0 illegal chars, 8 demerits           
    Score for Windows-1254: 0 illegal chars, 8 demerits
    Score for EUC-JP: 1 illegal chars, 4 demerits 
    Score for EUC-TW: 1 illegal chars, 4 demerits 
    Score for KOI8-R: 0 illegal chars, 8 demerits
    Score for BIG-5: 0 illegal chars, 36 demerits                                                   
 
    Score for ISO-2022-KR: 1 illegal chars, 4 demerits
    Score for ISO-2022-JP: 1 illegal chars, 4 demerits                                             
    Score for GB18030: 0 illegal chars, 36 demerits
    Score for UTF-32: 1 illegal chars, 0 demerits
    Score for UTF-32BE: 1 illegal chars, 0 demerits                                                 
 
    Score for UTF-32LE: 1 illegal chars, 0 demerits                                                 
 
    Score for UTF-16: 0 illegal chars, 91 demerits  
    Score for UTF-16BE: 0 illegal chars, 91 demerits
    Score for UTF-16LE: 0 illegal chars, 91 demerits                                                
 
    Score for UTF-7: 1 illegal chars, 4 demerits                                                    
 
    Score for UTF7-IMAP: 1 illegal chars, 4 demerits
    Score for ASCII: 1 illegal chars, 4 demerits
    Score for SJIS: 1 illegal chars, 4 demerits                                                     
 
    Score for eucJP-win: 1 illegal chars, 4 demerits                                                
 
    Score for EUC-JP-2004: 1 illegal chars, 4 demerits
    Score for SJIS-Mobile#DOCOMO: 0 illegal chars, 36 demerits
    Score for SJIS-Mobile#KDDI: 0 illegal chars, 36 demerits                                        

    Score for SJIS-Mobile#SOFTBANK: 0 illegal chars, 36 demerits
    Score for SJIS-mac: 1 illegal chars, 4 demerits
    Score for SJIS-2004: 0 illegal chars, 7 demerits
    Score for UTF-8-Mobile#DOCOMO: 1 illegal chars, 5 demerits
    Score for UTF-8-Mobile#KDDI-A: 1 illegal chars, 5 demerits
    Score for UTF-8-Mobile#KDDI-B: 1 illegal chars, 5 demerits
    Score for UTF-8-Mobile#SOFTBANK: 1 illegal chars, 5 demerits
    Score for CP932: 0 illegal chars, 36 demerits
    Score for CP51932: 1 illegal chars, 4 demerits
    Score for JIS: 1 illegal chars, 4 demerits
    Score for ISO-2022-JP-MS: 1 illegal chars, 4 demerits
    Score for Windows-1252: 0 illegal chars, 8 demerits
    Score for Windows-1254: 0 illegal chars, 8 demerits
    Score for EUC-CN: 1 illegal chars, 4 demerits
    Score for CP936: 0 illegal chars, 36 demerits
    Score for HZ: 1 illegal chars, 4 demerits
    Score for CP950: 0 illegal chars, 36 demerits
    Score for EUC-KR: 1 illegal chars, 4 demerits
    Score for UHC: 1 illegal chars, 4 demerits
    Score for Windows-1251: 0 illegal chars, 8 demerits
    Score for CP866: 0 illegal chars, 8 demerits
    Score for KOI8-U: 0 illegal chars, 8 demerits
    Score for ArmSCII-8: 0 illegal chars, 37 demerits 
    Score for CP850: 0 illegal chars, 8 demerits
    Score for ISO-2022-JP-2004: 1 illegal chars, 4 demerits
    Score for ISO-2022-JP-MOBILE#KDDI: 1 illegal chars, 4 demerits
    Score for CP50220: 1 illegal chars, 4 demerits
    Score for CP50221: 1 illegal chars, 4 demerits
    Score for CP50222: 1 illegal chars, 4 demerits
    SJIS-2004

Key lines are:

    Score for ISO-8859-1: 0 illegal chars, 8 demerits                                               
 
    Score for SJIS-2004: 0 illegal chars, 7 demerits

So what we have here is a case where the heuristics employed by mb_detect_encoding are not strong
enough to detect a significant difference in likelihood between ISO-8859-1 and SJIS-2004. SJIS-2004
happens to win out by a tiny margin, and we don't get the answer which was desired.

In ISO-8859-1 the string decodes to:

Iksiñski

And in SJIS-2004:

Iksi卧ki

It may look obvious that we wanted ñs and not 卧, but the current implementation of
mb_detect_encoding is based on inspecting codepoints one by one and seeing how many codepoints there
are which are 'rare' across all of the world's most common languages.
"ñ" and "s" are not rare (of course), and "卧" is also a fairly
common word in Chinese. So mb_detect_encoding can't see any difference between the two
decodings as far as rare codepoints go.

It also applies a small penalty to longer strings, which is necessary to avoid having *everything*
detected as a single-byte encoding where every possible byte value decodes to a codepoint which is
not rare.

Since there are no 'rare' codepoints in either decoding, and using SJIS-2004 results in a
slightly shorter output than ISO-8859-1, the function goes for SJIS-2004.

I'm trying to think of a way to tweak the heuristics to get the output which Alec wants on this
string, *without* making detection accuracy worse on a bunch of other possible inputs. It's
tricky. We can make it provide the desired answer on this particular example, but we may trash lots
and lots of other equally realistic  cases in the process.

I think the one thing we could do, which has not been done yet, is to look at *sequences* of
codepoints and judge them as likely or unlikely, rather than single codepoints. That has the
potential to significantly boost detection accuracy across the board, rather than on just one
cherry-picked example.

Of course, doing more checks will make the function a bit slower, which is a concern. We want it to
be as accurate as possible, but we also want it to be fast.

The bigger issue is where we would find the data to tell us which sequences of codepoints are likely
and which are unlikely. It would require gathering a big corpus of text in various languages which
we can analyze. And just a 'big' corpus doesn't guarantee that the results will be
good; there has to be enough data, but it also has to be balanced, good-quality data.

Then once we have that big corpus and can measure the frequency of various sequences of codepoints,
how much memory are we willing to give to the resulting tables? Right now I am using 8KB for a bit
vector (1 bit for each Unicode codepoint from U+0000 to U+FFFF). It would definitely take more than
that to get any useful results, but how much? I don't know. I anticipate that something like a
Bloom filter would be used to avoid consuming massive gobs of memory.

Maybe rather than looking at sequences of codepoints, we would look at sequences of codepoint
'types': "A Latin character, followed by punctuation, followed by whitespace,
followed by another Latin character..."

Not sure how much that would actually help us to boost accuracy. It would definitely reduce the size
of the needed corpus.

Anyways, if Nikita or someone else has smarter ideas than me, I would love to hear them. Or if
someone wants to help putting a good corpus together, I would be willing to write the code to use
it, but gathering the corpus is more work than I am ready to do now.

Thoughts?

------------------------------------------------------------------------
[2021-11-19 08:07:07] alec at alec dot pl

Nice to hear that, but we're running out of time. Any chance this bug can be fixed before the
final 8.1.0 release?

------------------------------------------------------------------------
[2021-11-13 21:20:12] alexdowad@php.net

Alec, you are nothing short of a fabulous tester. Well done.

I am going to investigate your latest finding... just not right now. At the moment, I am cooking
something up which will hopefully make almost every mbstring function several times faster. Soon to
be served up, hot and fresh, in a upcoming PHP release. Bon appetit!

------------------------------------------------------------------------
[2021-11-04 17:35:39] alec at alec dot pl

After testing rc5 I have one last case.

https://3v4l.org/9o7ft

There is a reasonably short input ('Iksi' . chr(241) . 'ski') recognized as
SJIS-2004, while I'd expect ISO-8859-2 or ISO-8859-1. It is also interesting that without the
second argument the result would be UTF-8.

I understand that such a short input may be confusing for the detector, but maybe there's still
some bug.

------------------------------------------------------------------------
[2021-11-01 19:40:34] patrickallaert@php.net

This is part of PHP 8.1.0RC5.

------------------------------------------------------------------------


The remainder of the comments for this report are too long. To view
the rest of the comments, please view the bug report online at

    https://bugs.php.net/bug.php?id=81390


--
Edit this bug report at https://bugs.php.net/bug.php?id=81390&edit=1


Thread (37 messages)

« previous php.bugs (#237932) next »