Bug #72933 [Opn]: mb_detect_encoding analyzing only the first byte of a string
| From: | yohgaki@php.net | Date: | Wed, 31 Aug 2016 05:24:31 +0000 |
| Subject: | Bug #72933 [Opn]: mb_detect_encoding analyzing only the first byte of a string | ||
| References: | 1 | Groups: | php.bugs |
| Request: | Send a blank email to php-bugs+get-203694@lists.php.net to get a copy of this message | ||
Edit report at https://bugs.php.net/bug.php?id=72933&edit=1
ID: 72933
Updated by: yohgaki@php.net
Reported by: paul dot crovella at gmail dot com
Summary: mb_detect_encoding analyzing only the first byte of
a string
Status: Open
Type: Bug
Package: mbstring related
PHP Version: Irrelevant
Block user comment: N
Private report: N
New Comment:
FYI. When you specify single encoding, the proper function for this task is
"mb_check_encoding()" rather than "mb_detect_encoding()".
None the less, mb_detect_encoding() should work correctly though.
Previous Comments:
------------------------------------------------------------------------
[2016-08-24 12:50:25] paul dot crovella at gmail dot com
Additional finding: it seems to be a problem when only a single encoding is given on the list to
detect. Even simply repeating UTF-8 achieves the expected result, e.g. mb_detect_encoding($str,
'UTF-8, UTF-8')
https://3v4l.org/g4BOv
------------------------------------------------------------------------
[2016-08-24 12:37:00] paul dot crovella at gmail dot com
Description:
------------
In non-strict mode it appears only the first byte of a string is being checked by mb_detect_encoding
in some circumstances.
For example, the byte 0xf8 is not allowed anywhere in UTF-8. When placed at the start of the string
mb_detect_encoding() properly returns false for it regardless of which mode is used. However if any
valid UTF-8 byte occurs at the beginning of the string mb_detect_encoding() in non-strict mode will
declare it as UTF-8.
This is also evident with a string like "\xe1\xe9\xf3\xfa", which is the ISO-8859-1
encoded version of "áéóú". The first byte, 0xe1, is allowed in UTF-8 as the
first of a multi-byte character - however the string as a whole is invalid and may not occur.
The suspect code is at:
https://github.com/php/php-src/blob/c72282a13b12b7e572469eba7a7ce593d900a8a2/ext/mbstring/libmbfl/mbfl/mbfilter.c#L746-L761
The problem exists in all current versions of PHP:
https://3v4l.org/b9b6q
Test script:
---------------
// This returns as expected.
$str = "\xf8foo";
var_dump(
mb_detect_encoding($str, 'UTF-8'), // bool(false)
mb_detect_encoding($str, 'UTF-8', true) // bool(false)
);
// This does not.
$str = "foo\xf8";
var_dump(
mb_detect_encoding($str, 'UTF-8'), // string(5) "UTF-8"
mb_detect_encoding($str, 'UTF-8', true) // bool(false)
);
// Nor does this.
$str = "\xe1\xe9\xf3\xfa";
var_dump(
mb_detect_encoding($str, 'UTF-8'), // string(5) "UTF-8"
mb_detect_encoding($str, 'UTF-8', true) // bool(false)
);
Expected result:
----------------
bool(false)
bool(false)
bool(false)
bool(false)
bool(false)
bool(false)
Actual result:
--------------
bool(false)
bool(false)
string(5) "UTF-8"
bool(false)
string(5) "UTF-8"
bool(false)
------------------------------------------------------------------------
--
Edit this bug report at https://bugs.php.net/bug.php?id=72933&edit=1