#15630 [Opn]: imap_utf7_decode appears to be broken
| From: | robert dot marchand at umontreal dot ca | Date: | Mon, 12 Aug 2002 19:47:29 +0000 |
| Subject: | #15630 [Opn]: imap_utf7_decode appears to be broken | ||
| References: | 1 | Groups: | php.bugs |
| Request: | Send a blank email to php-bugs+get-16552@lists.php.net to get a copy of this message | ||
ID: 15630
User updated by: robert.marchand@umontreal.ca
Reported By: robert.marchand@umontreal.ca
Status: Open
Bug Type: IMAP related
Operating System: SGI Irix 6.5
PHP Version: 4.2.2
New Comment:
Hi,
this is to confirm that there is a SGI specific issue because I
tested the new patch on a Linux Redhat 7.2 System with PHP 4.2.2
compiled manually. Here is the output from the test:
folder (modified UTF-7): &AMk-l&AOk-ments envoy&AOk-s
mb_convert_encoding test
folder decoded: [�léments envoyés]
(hexa: c3 89 6c c3 a9 6d 65 6e 74 73 20 65 6e 76 6f 79 c3 a9 73 )
encoded again: [&AMk-l&AOk-ments envoy&AOk-s]
decoded again: [�léments envoyés]
(hexa: c3 89 6c c3 a9 6d 65 6e 74 73 20 65 6e 76 6f 79 c3 a9 73 )
imap_utf7_decode test
folder decoded: [
(hexa: 0 c9 0 6c 0 e9 0 6d 0 65 0 6e 0 74 0 73 0 20 0 65 0 6e 0 76 0 6f
0 79 0 e9 0 73 )
encoded again: [&AMk-l&AOk-ments envoy&AOk-s]
decoded again: [
(hexa: 0 c9 0 6c 0 e9 0 6d 0 65 0 6e 0 74 0 73 0 20 0 65 0 6e 0 76 0 6f
0 79 0 e9 0 73 )
I have also try on an O2 SGI box (it is 32 bit) and it is the same as
all SGI boxes I have tried.
I'll take a look when I have a moment.
Thanks.
Previous Comments:
------------------------------------------------------------------------
[2002-08-12 13:57:01] robert.marchand@umontreal.ca
Hi,
you're write about the "utf8" thing. I was meaning "8bit". For the
rest, I cannot change the software I use (Horde/IMP) because it is not
me who wrote it. I can assure you it will break if you go with your
mods. This is for the general problem.
Now it seems I have a specific problem here with my SGI platform. Here
is what I get with your patch:
folder (modified UTF-7): test&AN9ZJw-
mb_convert_encoding test
folder decoded: [testÃY大]
(hexa: 74 65 73 74 c3 9f e5 a4 a7 )
encoded again: [test&AN9ZJw-]
decoded again: [testÃY大]
(hexa: 74 65 73 74 c3 9f e5 a4 a7 )
imap_utf7_decode test
folder decoded: [
(hexa: 0 74 0 65 0 73 0 74 0 d0 59 24 )
encoded again: [test&ANBZJA-]
decoded again: [
(hexa: 0 74 0 65 0 73 0 74 5 d9 59 24 )
Here is another sample:
folder (modified UTF-7): &AMk-l&AOk-ments envoy&AOk-s
mb_convert_encoding test
folder decoded: [�léments envoyés]
(hexa: c3 89 6c c3 a9 6d 65 6e 74 73 20 65 6e 76 6f 79 c3 a9 73 )
encoded again: [&AMk-l&AOk-ments envoy&AOk-s]
decoded again: [�léments envoyés]
(hexa: c3 89 6c c3 a9 6d 65 6e 74 73 20 65 6e 76 6f 79 c3 a9 73 )
imap_utf7_decode test
folder decoded: [ Ó
(hexa: 6 d3 0 6c 0 e0 0 6d 0 65 0 6e 0 74 0 73 0 20 0 65 0 6e 0 76 0 6f
0 79 0 e0 0 73 )
encoded again: [&BPp-l&A,g-ments envoy&Afa-s]
decoded again: [ û
(hexa: 4 fb 0 6c 0 fb 0 6d 0 65 0 6e 0 74 0 73 0 20 0 65 0 6e 0 76 0 6f
0 79 0 fc 0 73 )
Here is the PHP test page to generate this output:
<HTML>
<HEAD>
<TITLE>Test UTF7</TITLE>
<META HTTP-EQUIV="Content-Type" CONTENT="text/html;charset=utf-16">
</HEAD>
<BODY>
<?
function hexstr($s)
{
echo "(hexa: ";
for ($i=0;$i<strlen($s);$i++) {
echo dechex(ord($s[$i])), " ";
}
echo ")<br>";
}
//$folder = 'test&AN9ZJw-';
$folder = '&AMk-l&AOk-ments envoy&AOk-s';
echo "folder (modified UTF-7): $folder<BR><BR>\n";
echo "<strong>mb_convert_encoding test</strong><BR>\n";
$test = $folder;
$test = mb_convert_encoding($test, "UTF-8", "UTF7-IMAP");
echo " folder decoded: [$test]<BR>\n";
hexstr($test);
$test = mb_convert_encoding($test, "UTF7-IMAP", "UTF-8");
echo "encoded again: [", $test, "]<BR>\n";
$test = mb_convert_encoding($test, "UTF-8", "UTF7-IMAP");
echo "decoded again: [", $test, "]<BR>\n";
hexstr($test);
echo "<BR><strong>imap_utf7_decode test</strong><BR>\n";
$test = $folder;
$test = imap_utf7_decode($test);
echo "folder decoded: [", $test, "]<BR>\n";
hexstr($test);
$test = imap_utf7_encode($test);
echo "encoded again: [", $test, "]<BR>\n";
$test = imap_utf7_decode($test);
echo "decoded again: [", $test, "]<BR>\n";
hexstr($test);
?>
</BODY>
</HTML>
I am on a 64 bit platform. Could this be related to a wrapping shift?
There is definitly something wrong here.
Thanks.
------------------------------------------------------------------------
[2002-08-09 10:57:10] gamid@isayev.net
Robert Marchand wrote:
> this will not work without changing current applications.
Now it is not working at all for non-ASCII characters.
Example:
For IMAP folder name "test&WSc-" ("test" + chinese character) current
imap_utf7_decode() returns "testY'"
For IMAP folder name "testY'", current imap_utf7_decode() also returns
"testY'"
So, what you will do in this case?
> The problem is that these function try to encode and decode without
> knowing the charset used.
1) imap_utf7_decode() does not need to know charset of input string,
because input string is encoded in modified UTF7
2) if you specify charset for imap_utf7_decode() output string, what
will you do when IMAP folder name has characters from different
charsets (example: "test&BCQA31kn-" - ASCII, Russian, German,
Chinese)?
> As it is now, 8 bit is expected from imap_utf7_decode.
<...skiped...>
> It should really be:
> imap_utf7_utf8_decode
> imap_utf7_utf16_decode (patched version)
> imap_utf8_utf7_encode
> imap_utf16_utf7_encode (patched version)
I think you are confusing "8 bit" and UTF-8.
UTF-8 encoded character is "8 bit" only for ASCII characters. For
non-ASCII characters UTF-8 will be two and more bytes. So,
imap_utf7_decode() != imap_utf7_utf8_decode().
Gamid Isayev
------------------------------------------------------------------------
[2002-08-08 17:23:30] robert.marchand@umontreal.ca
Hi,
this will not work without changing current applications.
As it is now, 8 bit is expected from imap_utf7_decode. The problem is
that these function try to encode and decode without knowing the
charset used. It should really be:
imap_utf7_utf8_decode
imap_utf7_utf16_decode (patched version)
imap_utf8_utf7_encode
imap_utf16_utf7_encode (patched version)
Thanks.
------------------------------------------------------------------------
[2002-08-08 16:49:46] kalowsky@php.net
Since I have no way to test this, can anyone else confirm or deny that
this patch works? I'd rather not commit blindly.
------------------------------------------------------------------------
[2002-08-08 15:05:47] gamid@netilla.com
Robert,
The following patch fixes both imap_utf7_encode() and
imap_utf7_decode() to work with UTF-16.
PS: this patch is for PHP 4.2.2, the patch for CVS is posted in the
php.dev
Gamid Isayev
--- php_imap.c Wed Aug 7 15:45:53 2002
+++ php_imap.c Thu Aug 8 14:24:16 2002
@@ -2215,14 +2215,14 @@
php_error(E_WARNING, "imap_utf7_decode:
Invalid modified UTF-7 character: `%c'", *inp);
RETURN_FALSE;
} else if (*inp != '&') {
- outlen++;
+ outlen += 2;
} else if (inp + 1 == endp) {
php_error(E_WARNING, "imap_utf7_decode:
Unexpected end of string");
RETURN_FALSE;
} else if (inp[1] != '-') {
state = ST_DECODE0;
} else {
- outlen++;
+ outlen += 2;
inp++;
}
} else if (*inp == '-') {
@@ -2272,8 +2272,11 @@
if (*inp == '&' && inp[1] != '-') {
state = ST_DECODE0;
}
- else if ((*outp++ = *inp) == '&') {
- inp++;
+ else {
+ *outp++ = 0x00;
+ if ((*outp++ = *inp) == '&') {
+ inp++;
+ }
}
}
else if (*inp == '-') {
@@ -2349,29 +2352,42 @@
outlen = 0;
state = ST_NORMAL;
endp = (inp = in) + inlen;
- while (inp < endp) {
+ while (inp < endp || state != ST_NORMAL) {
if (state == ST_NORMAL) {
- if (SPECIAL(*inp)) {
+ if (*inp == 0x00 && *(inp+1) < 0x80) {
+ /* ASCII character */
+ outlen++; // for ASCII
char
+ if (*(inp+1) == '&')
+ outlen++; // for '-'
+ inp += 2;
+ } else {
+ /* begin encoding */
state = ST_ENCODE0;
- outlen++;
- } else if (*inp++ == '&') {
+ outlen++; // for '&'
+ }
+ } else if (inp == endp || (*inp == 0x00 && *(inp+1) <
0x80)) {
+ /* flush overflow and terminate region */
+ if (state != ST_ENCODE0) {
outlen++;
}
- outlen++;
- } else if (!SPECIAL(*inp)) {
+ outlen++; // for '-'
state = ST_NORMAL;
} else {
- /* ST_ENCODE0 -> ST_ENCODE1 - two chars
- * ST_ENCODE1 -> ST_ENCODE2 - one char
- * ST_ENCODE2 -> ST_ENCODE0 - one char
- */
- if (state == ST_ENCODE2) {
- state = ST_ENCODE0;
- }
- else if (state++ == ST_ENCODE0) {
- outlen++;
+ switch (state) {
+ case ST_ENCODE0:
+ outlen++;
+ state = ST_ENCODE1;
+ break;
+ case ST_ENCODE1:
+ outlen++;
+ state = ST_ENCODE2;
+ break;
+ case ST_ENCODE2:
+ outlen += 2;
+ state = ST_ENCODE0;
+ case ST_NORMAL:
+ break;
}
- outlen++;
inp++;
}
}
@@ -2388,14 +2404,17 @@
endp = (inp = in) + inlen;
while (inp < endp || state != ST_NORMAL) {
if (state == ST_NORMAL) {
- if (SPECIAL(*inp)) {
+ if (*inp == 0x00 && *(inp+1) < 0x80) {
+ /* ASCII character */
+ inp++;
+ if ((*outp++ = *inp++) == '&')
+ *outp++ = '-';
+ } else {
/* begin encoding */
*outp++ = '&';
state = ST_ENCODE0;
- } else if ((*outp++ = *inp++) == '&') {
- *outp++ = '-';
}
- } else if (inp == endp || !SPECIAL(*inp)) {
+ } else if (inp == endp || (*inp == 0x00 && *(inp+1) <
0x80)) {
/* flush overflow and terminate region */
if (state != ST_ENCODE0) {
*outp++ = B64(*outp);
------------------------------------------------------------------------
The remainder of the comments for this report are too long. To view
the rest of the comments, please view the bug report online at
http://bugs.php.net/15630
--
Edit this bug report at http://bugs.php.net/?id=15630&edit=1