Bug #45993 [Com]: mb_detect_encoding and mb_check_encoding results are dissonant

From: Date: Fri, 11 Dec 2015 17:40:35 +0000
Subject: Bug #45993 [Com]: mb_detect_encoding and mb_check_encoding results are dissonant
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-197805@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=45993&edit=1 ID: 45993 Comment by: nathanael at gnat dot ca Reported by: mtrojan at transline dot de Summary: mb_detect_encoding and mb_check_encoding results are dissonant Status: Open Type: Bug Package: mbstring related Operating System: Windows XP PHP Version: 5.2.6 Block user comment: N Private report: N New Comment: So there seems to be some regression here between 5.5 and 5.6. I have a unit test for a project. It took a UTF-16 encoded file (with BOM), copied to a tmp dir, then detects encoding and requests mb_convert_encoding($fileContent,'UTF-8') the file. On php 5.5 the file is converted to UTF-8 properly. On 5.6 (on linux and windows) and 7 (linux) it fails. The BOM becomes ?? and then the file is detected as ASCII. Using if ($content[0]==chr(0xff) && $content[1]==chr(0xfe)) { echo 'UTF-16'; } does detect it as UTF-16, but I'd like to be able to detect the files that are multi-byte and convert them to UTF-8. Previous Comments: ------------------------------------------------------------------------ [2014-04-04 14:32:42] soapergem at gmail dot com I came here to report essentially this same bug. In fact I think this bug is directly related to bugs 51563, 64667, 63433, and even 38138. I have a UTF-16LE encoded CSV file and fgetcsv() was failing on it. So I got to learn all about different character encodings today! I read on another bug report from one of the PHP devs that the purpose of mb_detect_encoding() is to "detect which multibyte encoding is in use." It's failing at that right now. When I run mb_detect_encoding() on a UTF-16 encoded string, either it says ASCII (if it is not the first line of the file), or it just returns FALSE (if it is the first line, which includes the BOM). On the other hand, if I run mb_check_encoding($str, 'UTF-16') then it seems I get TRUE for all except the first line. I'm using PHP 5.5.10 by the way. ------------------------------------------------------------------------ [2012-01-02 04:22:49] Apollo880 at gmail dot com Bug with correct encoding detection. function detect_enc($str) { $awe = mb_list_encodings(); unset($awe[0], $awe[1], $awe[2]); foreach ($awe as $enctype) { if (mb_check_encoding($str, $enctype) === true) return $enctype; } return false; } echo detect_enc('String_encoded_to_Windows-1251'); // Return 'byte2be'. It's a fail. ------------------------------------------------------------------------ [2008-11-10 07:30:32] mtrojan at transline dot de Of course, comparing the beginning of a file with the UTF-16 BOM can be used to detect UTF-16 encoding. But what do you do with UTF-16 encoded files where no BOM is set? ------------------------------------------------------------------------ [2008-11-08 02:20:46] hirokawa@php.net mb_detect_encoding does not support the UTF-16/UTF-16BE encoding detection. Because UTF-16 isn't byte stream encoding like UTF-8, we cannot detect the encoding as other byte stream encoding. The file encoded in UTF-16 can be detected easily using BOM, it is like, if ($content[0]==chr(0xff) && $content[1]==chr(0xfe)) { echo 'UTF-16'; } else if ($content[0]==chr(0xfe) && $content[1]==chr(0xff)) { echo 'UTF-16BE'; } ------------------------------------------------------------------------ [2008-10-26 23:01:49] jani@php.net Assigned to the mbstring maintainer. ------------------------------------------------------------------------ The remainder of the comments for this report are too long. To view the rest of the comments, please view the bug report online at https://bugs.php.net/bug.php?id=45993 -- Edit this bug report at https://bugs.php.net/bug.php?id=45993&edit=1

« previous php.bugs (#197805) next »