Bug #72933 [Opn]: mb_detect_encoding analyzing only the first byte of a string

From: Date: Wed, 31 Aug 2016 05:24:31 +0000
Subject: Bug #72933 [Opn]: mb_detect_encoding analyzing only the first byte of a string
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-203694@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=72933&edit=1 ID: 72933 Updated by: yohgaki@php.net Reported by: paul dot crovella at gmail dot com Summary: mb_detect_encoding analyzing only the first byte of a string Status: Open Type: Bug Package: mbstring related PHP Version: Irrelevant Block user comment: N Private report: N New Comment: FYI. When you specify single encoding, the proper function for this task is "mb_check_encoding()" rather than "mb_detect_encoding()". None the less, mb_detect_encoding() should work correctly though. Previous Comments: ------------------------------------------------------------------------ [2016-08-24 12:50:25] paul dot crovella at gmail dot com Additional finding: it seems to be a problem when only a single encoding is given on the list to detect. Even simply repeating UTF-8 achieves the expected result, e.g. mb_detect_encoding($str, 'UTF-8, UTF-8') https://3v4l.org/g4BOv ------------------------------------------------------------------------ [2016-08-24 12:37:00] paul dot crovella at gmail dot com Description: ------------ In non-strict mode it appears only the first byte of a string is being checked by mb_detect_encoding in some circumstances. For example, the byte 0xf8 is not allowed anywhere in UTF-8. When placed at the start of the string mb_detect_encoding() properly returns false for it regardless of which mode is used. However if any valid UTF-8 byte occurs at the beginning of the string mb_detect_encoding() in non-strict mode will declare it as UTF-8. This is also evident with a string like "\xe1\xe9\xf3\xfa", which is the ISO-8859-1 encoded version of "áéóú". The first byte, 0xe1, is allowed in UTF-8 as the first of a multi-byte character - however the string as a whole is invalid and may not occur. The suspect code is at: https://github.com/php/php-src/blob/c72282a13b12b7e572469eba7a7ce593d900a8a2/ext/mbstring/libmbfl/mbfl/mbfilter.c#L746-L761 The problem exists in all current versions of PHP: https://3v4l.org/b9b6q Test script: --------------- // This returns as expected. $str = "\xf8foo"; var_dump( mb_detect_encoding($str, 'UTF-8'), // bool(false) mb_detect_encoding($str, 'UTF-8', true) // bool(false) ); // This does not. $str = "foo\xf8"; var_dump( mb_detect_encoding($str, 'UTF-8'), // string(5) "UTF-8" mb_detect_encoding($str, 'UTF-8', true) // bool(false) ); // Nor does this. $str = "\xe1\xe9\xf3\xfa"; var_dump( mb_detect_encoding($str, 'UTF-8'), // string(5) "UTF-8" mb_detect_encoding($str, 'UTF-8', true) // bool(false) ); Expected result: ---------------- bool(false) bool(false) bool(false) bool(false) bool(false) bool(false) Actual result: -------------- bool(false) bool(false) string(5) "UTF-8" bool(false) string(5) "UTF-8" bool(false) ------------------------------------------------------------------------ -- Edit this bug report at https://bugs.php.net/bug.php?id=72933&edit=1

« previous php.bugs (#203694) next »