Bug #45993 [Com]: mb_detect_encoding and mb_check_encoding results are dissonant
| From: | nathanael at gnat dot ca | Date: | Fri, 11 Dec 2015 17:40:35 +0000 |
| Subject: | Bug #45993 [Com]: mb_detect_encoding and mb_check_encoding results are dissonant | ||
| References: | 1 | Groups: | php.bugs |
| Request: | Send a blank email to php-bugs+get-197805@lists.php.net to get a copy of this message | ||
Edit report at https://bugs.php.net/bug.php?id=45993&edit=1
ID: 45993
Comment by: nathanael at gnat dot ca
Reported by: mtrojan at transline dot de
Summary: mb_detect_encoding and mb_check_encoding results are
dissonant
Status: Open
Type: Bug
Package: mbstring related
Operating System: Windows XP
PHP Version: 5.2.6
Block user comment: N
Private report: N
New Comment:
So there seems to be some regression here between 5.5 and 5.6. I have a unit test for a project. It
took a UTF-16 encoded file (with BOM), copied to a tmp dir, then detects encoding and requests
mb_convert_encoding($fileContent,'UTF-8') the file. On php 5.5 the file is converted to
UTF-8 properly. On 5.6 (on linux and windows) and 7 (linux) it fails. The BOM becomes ?? and then
the file is detected as ASCII.
Using
if ($content[0]==chr(0xff) && $content[1]==chr(0xfe)) {
echo 'UTF-16';
}
does detect it as UTF-16, but I'd like to be able to detect the files that are multi-byte and
convert them to UTF-8.
Previous Comments:
------------------------------------------------------------------------
[2014-04-04 14:32:42] soapergem at gmail dot com
I came here to report essentially this same bug. In fact I think this bug is directly related to
bugs 51563, 64667, 63433, and even 38138.
I have a UTF-16LE encoded CSV file and fgetcsv() was failing on it. So I got to learn all about
different character encodings today!
I read on another bug report from one of the PHP devs that the purpose of mb_detect_encoding() is to
"detect which multibyte encoding is in use." It's failing at that right now. When I
run mb_detect_encoding() on a UTF-16 encoded string, either it says ASCII (if it is not the first
line of the file), or it just returns FALSE (if it is the first line, which includes the BOM). On
the other hand, if I run mb_check_encoding($str, 'UTF-16') then it seems I get TRUE for
all except the first line.
I'm using PHP 5.5.10 by the way.
------------------------------------------------------------------------
[2012-01-02 04:22:49] Apollo880 at gmail dot com
Bug with correct encoding detection.
function detect_enc($str)
{
$awe = mb_list_encodings();
unset($awe[0], $awe[1], $awe[2]);
foreach ($awe as $enctype)
{
if (mb_check_encoding($str, $enctype) === true) return $enctype;
}
return false;
}
echo detect_enc('String_encoded_to_Windows-1251'); // Return 'byte2be'.
It's a fail.
------------------------------------------------------------------------
[2008-11-10 07:30:32] mtrojan at transline dot de
Of course, comparing the beginning of a file with the UTF-16 BOM can be used to detect UTF-16
encoding. But what do you do with UTF-16 encoded files where no BOM is set?
------------------------------------------------------------------------
[2008-11-08 02:20:46] hirokawa@php.net
mb_detect_encoding does not support the UTF-16/UTF-16BE
encoding detection. Because UTF-16 isn't byte stream encoding like UTF-8, we cannot detect the
encoding as other byte stream encoding.
The file encoded in UTF-16 can be detected easily using BOM,
it is like,
if ($content[0]==chr(0xff) && $content[1]==chr(0xfe)) {
echo 'UTF-16';
} else if ($content[0]==chr(0xfe) && $content[1]==chr(0xff)) {
echo 'UTF-16BE';
}
------------------------------------------------------------------------
[2008-10-26 23:01:49] jani@php.net
Assigned to the mbstring maintainer.
------------------------------------------------------------------------
The remainder of the comments for this report are too long. To view
the rest of the comments, please view the bug report online at
https://bugs.php.net/bug.php?id=45993
--
Edit this bug report at https://bugs.php.net/bug.php?id=45993&edit=1