Re: #15630 [Com]: imap_utf7_decode appears to be broken
| From: | Gamid Isayev | Date: | Wed, 07 Aug 2002 20:33:03 +0000 |
| Subject: | Re: #15630 [Com]: imap_utf7_decode appears to be broken | ||
| References: | 1 2 | Groups: | php.dev |
| Request: | Send a blank email to php-dev+get-86613@lists.php.net to get a copy of this message | ||
Robert Marchand wrote:
here is what I obtain with the patch applied (php-4.2.2):
Maybe there is something wrong with my installation here, but even then, I'm not sure adding zero bytes is the good solution. In my view it will break every application that now use these functions. At least this patch should be used only when --enable-mbstring is used.The following patch fixes both imap_utf7_encode() and imap_utf7_decode() to work with UTF-16. PS: this patch is for PHP 4.2.2, the patch for CVS is posted in the php.dev Gamid Isayev. --- php_imap.c Wed Aug 7 15:45:53 2002 +++ php_imap.c Wed Aug 7 15:54:37 2002 @@ -2215,14 +2215,14 @@
php_error(E_WARNING, "imap_utf7_decode: Invalid modified UTF-7 character: `%c'", *inp);
RETURN_FALSE;
} else if (*inp != '&') {
- outlen++;
+ outlen += 2;
} else if (inp + 1 == endp) {
php_error(E_WARNING, "imap_utf7_decode: Unexpected end of string");
RETURN_FALSE;
} else if (inp[1] != '-') {
state = ST_DECODE0;
} else {
- outlen++;
+ outlen += 2;
inp++;
}
} else if (*inp == '-') {
@@ -2272,8 +2272,11 @@
if (*inp == '&' && inp[1] != '-') {
state = ST_DECODE0;
}
- else if ((*outp++ = *inp) == '&') {
- inp++;
+ else {
+ *outp++ = 0x00;
+ if ((*outp++ = *inp) == '&') {
+ inp++;
+ }
}
}
else if (*inp == '-') {
@@ -2351,15 +2354,21 @@
endp = (inp = in) + inlen;
while (inp < endp) {
if (state == ST_NORMAL) {
- if (SPECIAL(*inp)) {
+ if (*inp == 0x00 && *(inp+1) < 0x80) {
+ /* ASCII character */
+ if (*(inp+1) == '&')
+ outlen++; // for '-'
+ inp += 2;
+ } else {
+ /* begin encoding */
state = ST_ENCODE0;
- outlen++;
- } else if (*inp++ == '&') {
- outlen++;
+ outlen++; // for '&'
}
outlen++;
- } else if (!SPECIAL(*inp)) {
+ } else if (*inp == 0x00 && *(inp+1) < 0x80) {
state = ST_NORMAL;
+ if (*(inp+1) != '&')
+ outlen++; // for '-'
} else {
/* ST_ENCODE0 -> ST_ENCODE1 - two chars
* ST_ENCODE1 -> ST_ENCODE2 - one char
@@ -2371,8 +2380,8 @@
else if (state++ == ST_ENCODE0) {
outlen++;
}
- outlen++;
- inp++;
+ outlen += 2;
+ inp += 2;
}
}
@@ -2388,14 +2397,17 @@
endp = (inp = in) + inlen;
while (inp < endp || state != ST_NORMAL) {
if (state == ST_NORMAL) {
- if (SPECIAL(*inp)) {
+ if (*inp == 0x00 && *(inp+1) < 0x80) {
+ /* ASCII character */
+ inp++;
+ if ((*outp++ = *inp++) == '&')
+ *outp++ = '-';
+ } else {
/* begin encoding */
*outp++ = '&';
state = ST_ENCODE0;
- } else if ((*outp++ = *inp++) == '&') {
- *outp++ = '-';
}
- } else if (inp == endp || !SPECIAL(*inp)) {
+ } else if (inp == endp || (*inp == 0x00 && *(inp+1) < 0x80)) {
/* flush overflow and terminate region */
if (state != ST_ENCODE0) {
*outp++ = B64(*outp);