Bug->Req #45993 [Opn]: mb_detect_encoding should support UTF-16

From: Date: Sun, 31 Jul 2016 13:56:34 +0000
Subject: Bug->Req #45993 [Opn]: mb_detect_encoding should support UTF-16
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-202776@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=45993&edit=1 ID: 45993 Updated by: cmb@php.net Reported by: mtrojan at transline dot de -Summary: mb_detect_encoding and mb_check_encoding results are dissonant +Summary: mb_detect_encoding should support UTF-16 Status: Open -Type: Bug +Type: Feature/Change Request Package: mbstring related Operating System: Windows XP PHP Version: 5.2.6 Block user comment: N Private report: N New Comment: > mb_detect_encoding does not seem to recognize UTF-16 encoded > files properly. That is expected behavior, that's already documented[1]: | For UTF-16, UTF-32, UCS2 and UCS4, encoding detection will fail | always. I'm therefore chaning to feature request. > The file encoded in UTF-16 can be detected easily using BOM, Albeit not reliably, because 0xFE and 0xFF are valid ISO-8859-* characters, for instance. Furthermore, a BOM is optional for UTF-16. @nathanael at gnat dot ca: that would be a different issue, so please open a separate ticket. [1] <http://php.net/manual/en/function.mb-detect-order.php> Previous Comments: ------------------------------------------------------------------------ [2015-12-11 17:40:34] nathanael at gnat dot ca So there seems to be some regression here between 5.5 and 5.6. I have a unit test for a project. It took a UTF-16 encoded file (with BOM), copied to a tmp dir, then detects encoding and requests mb_convert_encoding($fileContent,'UTF-8') the file. On php 5.5 the file is converted to UTF-8 properly. On 5.6 (on linux and windows) and 7 (linux) it fails. The BOM becomes ?? and then the file is detected as ASCII. Using if ($content[0]==chr(0xff) && $content[1]==chr(0xfe)) { echo 'UTF-16'; } does detect it as UTF-16, but I'd like to be able to detect the files that are multi-byte and convert them to UTF-8. ------------------------------------------------------------------------ [2014-04-04 14:32:42] soapergem at gmail dot com I came here to report essentially this same bug. In fact I think this bug is directly related to bugs 51563, 64667, 63433, and even 38138. I have a UTF-16LE encoded CSV file and fgetcsv() was failing on it. So I got to learn all about different character encodings today! I read on another bug report from one of the PHP devs that the purpose of mb_detect_encoding() is to "detect which multibyte encoding is in use." It's failing at that right now. When I run mb_detect_encoding() on a UTF-16 encoded string, either it says ASCII (if it is not the first line of the file), or it just returns FALSE (if it is the first line, which includes the BOM). On the other hand, if I run mb_check_encoding($str, 'UTF-16') then it seems I get TRUE for all except the first line. I'm using PHP 5.5.10 by the way. ------------------------------------------------------------------------ [2012-01-02 04:22:49] Apollo880 at gmail dot com Bug with correct encoding detection. function detect_enc($str) { $awe = mb_list_encodings(); unset($awe[0], $awe[1], $awe[2]); foreach ($awe as $enctype) { if (mb_check_encoding($str, $enctype) === true) return $enctype; } return false; } echo detect_enc('String_encoded_to_Windows-1251'); // Return 'byte2be'. It's a fail. ------------------------------------------------------------------------ [2008-11-10 07:30:32] mtrojan at transline dot de Of course, comparing the beginning of a file with the UTF-16 BOM can be used to detect UTF-16 encoding. But what do you do with UTF-16 encoded files where no BOM is set? ------------------------------------------------------------------------ [2008-11-08 02:20:46] hirokawa@php.net mb_detect_encoding does not support the UTF-16/UTF-16BE encoding detection. Because UTF-16 isn't byte stream encoding like UTF-8, we cannot detect the encoding as other byte stream encoding. The file encoded in UTF-16 can be detected easily using BOM, it is like, if ($content[0]==chr(0xff) && $content[1]==chr(0xfe)) { echo 'UTF-16'; } else if ($content[0]==chr(0xfe) && $content[1]==chr(0xff)) { echo 'UTF-16BE'; } ------------------------------------------------------------------------ The remainder of the comments for this report are too long. To view the rest of the comments, please view the bug report online at https://bugs.php.net/bug.php?id=45993 -- Edit this bug report at https://bugs.php.net/bug.php?id=45993&edit=1

« previous php.bugs (#202776) next »