#15630 [Opn]: imap_utf7_decode appears to be broken

From: Date: Mon, 12 Aug 2002 19:47:29 +0000
Subject: #15630 [Opn]: imap_utf7_decode appears to be broken
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-16552@lists.php.net to get a copy of this message
ID: 15630 User updated by: robert.marchand@umontreal.ca Reported By: robert.marchand@umontreal.ca Status: Open Bug Type: IMAP related Operating System: SGI Irix 6.5 PHP Version: 4.2.2 New Comment: Hi, this is to confirm that there is a SGI specific issue because I tested the new patch on a Linux Redhat 7.2 System with PHP 4.2.2 compiled manually. Here is the output from the test: folder (modified UTF-7): &AMk-l&AOk-ments envoy&AOk-s mb_convert_encoding test folder decoded: [Ã?léments envoyés] (hexa: c3 89 6c c3 a9 6d 65 6e 74 73 20 65 6e 76 6f 79 c3 a9 73 ) encoded again: [&AMk-l&AOk-ments envoy&AOk-s] decoded again: [Ã?léments envoyés] (hexa: c3 89 6c c3 a9 6d 65 6e 74 73 20 65 6e 76 6f 79 c3 a9 73 ) imap_utf7_decode test folder decoded: [ (hexa: 0 c9 0 6c 0 e9 0 6d 0 65 0 6e 0 74 0 73 0 20 0 65 0 6e 0 76 0 6f 0 79 0 e9 0 73 ) encoded again: [&AMk-l&AOk-ments envoy&AOk-s] decoded again: [ (hexa: 0 c9 0 6c 0 e9 0 6d 0 65 0 6e 0 74 0 73 0 20 0 65 0 6e 0 76 0 6f 0 79 0 e9 0 73 ) I have also try on an O2 SGI box (it is 32 bit) and it is the same as all SGI boxes I have tried. I'll take a look when I have a moment. Thanks. Previous Comments: ------------------------------------------------------------------------ [2002-08-12 13:57:01] robert.marchand@umontreal.ca Hi, you're write about the "utf8" thing. I was meaning "8bit". For the rest, I cannot change the software I use (Horde/IMP) because it is not me who wrote it. I can assure you it will break if you go with your mods. This is for the general problem. Now it seems I have a specific problem here with my SGI platform. Here is what I get with your patch: folder (modified UTF-7): test&AN9ZJw- mb_convert_encoding test folder decoded: [testÃY大] (hexa: 74 65 73 74 c3 9f e5 a4 a7 ) encoded again: [test&AN9ZJw-] decoded again: [testÃY大] (hexa: 74 65 73 74 c3 9f e5 a4 a7 ) imap_utf7_decode test folder decoded: [ (hexa: 0 74 0 65 0 73 0 74 0 d0 59 24 ) encoded again: [test&ANBZJA-] decoded again: [ (hexa: 0 74 0 65 0 73 0 74 5 d9 59 24 ) Here is another sample: folder (modified UTF-7): &AMk-l&AOk-ments envoy&AOk-s mb_convert_encoding test folder decoded: [Ã?léments envoyés] (hexa: c3 89 6c c3 a9 6d 65 6e 74 73 20 65 6e 76 6f 79 c3 a9 73 ) encoded again: [&AMk-l&AOk-ments envoy&AOk-s] decoded again: [Ã?léments envoyés] (hexa: c3 89 6c c3 a9 6d 65 6e 74 73 20 65 6e 76 6f 79 c3 a9 73 ) imap_utf7_decode test folder decoded: [ Ó (hexa: 6 d3 0 6c 0 e0 0 6d 0 65 0 6e 0 74 0 73 0 20 0 65 0 6e 0 76 0 6f 0 79 0 e0 0 73 ) encoded again: [&BPp-l&A,g-ments envoy&Afa-s] decoded again: [ û (hexa: 4 fb 0 6c 0 fb 0 6d 0 65 0 6e 0 74 0 73 0 20 0 65 0 6e 0 76 0 6f 0 79 0 fc 0 73 ) Here is the PHP test page to generate this output: <HTML> <HEAD> <TITLE>Test UTF7</TITLE> <META HTTP-EQUIV="Content-Type" CONTENT="text/html;charset=utf-16"> </HEAD> <BODY> <? function hexstr($s) { echo "(hexa: "; for ($i=0;$i<strlen($s);$i++) { echo dechex(ord($s[$i])), " "; } echo ")<br>"; } //$folder = 'test&AN9ZJw-'; $folder = '&AMk-l&AOk-ments envoy&AOk-s'; echo "folder (modified UTF-7): $folder<BR><BR>\n"; echo "<strong>mb_convert_encoding test</strong><BR>\n"; $test = $folder; $test = mb_convert_encoding($test, "UTF-8", "UTF7-IMAP"); echo " folder decoded: [$test]<BR>\n"; hexstr($test); $test = mb_convert_encoding($test, "UTF7-IMAP", "UTF-8"); echo "encoded again: [", $test, "]<BR>\n"; $test = mb_convert_encoding($test, "UTF-8", "UTF7-IMAP"); echo "decoded again: [", $test, "]<BR>\n"; hexstr($test); echo "<BR><strong>imap_utf7_decode test</strong><BR>\n"; $test = $folder; $test = imap_utf7_decode($test); echo "folder decoded: [", $test, "]<BR>\n"; hexstr($test); $test = imap_utf7_encode($test); echo "encoded again: [", $test, "]<BR>\n"; $test = imap_utf7_decode($test); echo "decoded again: [", $test, "]<BR>\n"; hexstr($test); ?> </BODY> </HTML> I am on a 64 bit platform. Could this be related to a wrapping shift? There is definitly something wrong here. Thanks. ------------------------------------------------------------------------ [2002-08-09 10:57:10] gamid@isayev.net Robert Marchand wrote: > this will not work without changing current applications. Now it is not working at all for non-ASCII characters. Example: For IMAP folder name "test&WSc-" ("test" + chinese character) current imap_utf7_decode() returns "testY'" For IMAP folder name "testY'", current imap_utf7_decode() also returns "testY'" So, what you will do in this case? > The problem is that these function try to encode and decode without > knowing the charset used. 1) imap_utf7_decode() does not need to know charset of input string, because input string is encoded in modified UTF7 2) if you specify charset for imap_utf7_decode() output string, what will you do when IMAP folder name has characters from different charsets (example: "test&BCQA31kn-" - ASCII, Russian, German, Chinese)? > As it is now, 8 bit is expected from imap_utf7_decode. <...skiped...> > It should really be: > imap_utf7_utf8_decode > imap_utf7_utf16_decode (patched version) > imap_utf8_utf7_encode > imap_utf16_utf7_encode (patched version) I think you are confusing "8 bit" and UTF-8. UTF-8 encoded character is "8 bit" only for ASCII characters. For non-ASCII characters UTF-8 will be two and more bytes. So, imap_utf7_decode() != imap_utf7_utf8_decode(). Gamid Isayev ------------------------------------------------------------------------ [2002-08-08 17:23:30] robert.marchand@umontreal.ca Hi, this will not work without changing current applications. As it is now, 8 bit is expected from imap_utf7_decode. The problem is that these function try to encode and decode without knowing the charset used. It should really be: imap_utf7_utf8_decode imap_utf7_utf16_decode (patched version) imap_utf8_utf7_encode imap_utf16_utf7_encode (patched version) Thanks. ------------------------------------------------------------------------ [2002-08-08 16:49:46] kalowsky@php.net Since I have no way to test this, can anyone else confirm or deny that this patch works? I'd rather not commit blindly. ------------------------------------------------------------------------ [2002-08-08 15:05:47] gamid@netilla.com Robert, The following patch fixes both imap_utf7_encode() and imap_utf7_decode() to work with UTF-16. PS: this patch is for PHP 4.2.2, the patch for CVS is posted in the php.dev Gamid Isayev --- php_imap.c Wed Aug 7 15:45:53 2002 +++ php_imap.c Thu Aug 8 14:24:16 2002 @@ -2215,14 +2215,14 @@ php_error(E_WARNING, "imap_utf7_decode: Invalid modified UTF-7 character: `%c'", *inp); RETURN_FALSE; } else if (*inp != '&') { - outlen++; + outlen += 2; } else if (inp + 1 == endp) { php_error(E_WARNING, "imap_utf7_decode: Unexpected end of string"); RETURN_FALSE; } else if (inp[1] != '-') { state = ST_DECODE0; } else { - outlen++; + outlen += 2; inp++; } } else if (*inp == '-') { @@ -2272,8 +2272,11 @@ if (*inp == '&' && inp[1] != '-') { state = ST_DECODE0; } - else if ((*outp++ = *inp) == '&') { - inp++; + else { + *outp++ = 0x00; + if ((*outp++ = *inp) == '&') { + inp++; + } } } else if (*inp == '-') { @@ -2349,29 +2352,42 @@ outlen = 0; state = ST_NORMAL; endp = (inp = in) + inlen; - while (inp < endp) { + while (inp < endp || state != ST_NORMAL) { if (state == ST_NORMAL) { - if (SPECIAL(*inp)) { + if (*inp == 0x00 && *(inp+1) < 0x80) { + /* ASCII character */ + outlen++; // for ASCII char + if (*(inp+1) == '&') + outlen++; // for '-' + inp += 2; + } else { + /* begin encoding */ state = ST_ENCODE0; - outlen++; - } else if (*inp++ == '&') { + outlen++; // for '&' + } + } else if (inp == endp || (*inp == 0x00 && *(inp+1) < 0x80)) { + /* flush overflow and terminate region */ + if (state != ST_ENCODE0) { outlen++; } - outlen++; - } else if (!SPECIAL(*inp)) { + outlen++; // for '-' state = ST_NORMAL; } else { - /* ST_ENCODE0 -> ST_ENCODE1 - two chars - * ST_ENCODE1 -> ST_ENCODE2 - one char - * ST_ENCODE2 -> ST_ENCODE0 - one char - */ - if (state == ST_ENCODE2) { - state = ST_ENCODE0; - } - else if (state++ == ST_ENCODE0) { - outlen++; + switch (state) { + case ST_ENCODE0: + outlen++; + state = ST_ENCODE1; + break; + case ST_ENCODE1: + outlen++; + state = ST_ENCODE2; + break; + case ST_ENCODE2: + outlen += 2; + state = ST_ENCODE0; + case ST_NORMAL: + break; } - outlen++; inp++; } } @@ -2388,14 +2404,17 @@ endp = (inp = in) + inlen; while (inp < endp || state != ST_NORMAL) { if (state == ST_NORMAL) { - if (SPECIAL(*inp)) { + if (*inp == 0x00 && *(inp+1) < 0x80) { + /* ASCII character */ + inp++; + if ((*outp++ = *inp++) == '&') + *outp++ = '-'; + } else { /* begin encoding */ *outp++ = '&'; state = ST_ENCODE0; - } else if ((*outp++ = *inp++) == '&') { - *outp++ = '-'; } - } else if (inp == endp || !SPECIAL(*inp)) { + } else if (inp == endp || (*inp == 0x00 && *(inp+1) < 0x80)) { /* flush overflow and terminate region */ if (state != ST_ENCODE0) { *outp++ = B64(*outp); ------------------------------------------------------------------------ The remainder of the comments for this report are too long. To view the rest of the comments, please view the bug report online at http://bugs.php.net/15630 -- Edit this bug report at http://bugs.php.net/?id=15630&edit=1

« previous php.bugs (#16552) next »