Bug #66828 [Ana->Csd]: iconv_mime_encode Q-encoding longer than it should be
| From: | cmb@php.net | Date: | Sat, 22 Sep 2018 14:09:11 +0000 |
| Subject: | Bug #66828 [Ana->Csd]: iconv_mime_encode Q-encoding longer than it should be | ||
| References: | 1 | Groups: | php.bugs |
| Request: | Send a blank email to php-bugs+get-217194@lists.php.net to get a copy of this message | ||
Edit report at https://bugs.php.net/bug.php?id=66828&edit=1
ID: 66828
Updated by: cmb@php.net
Reported by: st_9876543210 at yahoo dot de
Summary: iconv_mime_encode Q-encoding longer than it should
be
-Status: Analyzed
+Status: Closed
Type: Bug
Package: ICONV related
PHP Version: Irrelevant
Block user comment: N
Private report: N
New Comment:
Automatic comment on behalf of cmbecker69@gmx.de
Revision: http://git.php.net/?p=php-src.git;a=commit;h=9cbe1283f70699b82ca4225705d40fcc73633dfb
Log: Fix #66828: iconv_mime_encode Q-encoding longer than it should be
Previous Comments:
------------------------------------------------------------------------
[2018-09-04 14:18:26] cmb@php.net
<https://github.com/php/php-src/pull/3492> is
supposed to fix this
issue.
------------------------------------------------------------------------
[2018-09-04 13:31:38] cmb@php.net
> It seems to me that the most appropriate solution would be to
> convert the complete input to the desired output charset, and to
> apply the encoding afterwards.
No, that can't work, since RFC 2047 specifies[1]:
| Each 'encoded-word' MUST represent an integral number of
| characters. A multi-octet character may not be split across
| adjacent 'encoded-word's.
> Simply removing the
/ 3[2] causes bug48289.phpt[3] to hang.
Indeed. That is because the out_size isn't changed[2] in this
case, since the +1 is not enough to enforce the division by 3 to
be greater than zero; we need to add 2 instead.
[1] <https://tools.ietf.org/html/rfc2047#page-8>
[2] <https://github.com/php/php-src/blob/php-7.3.0beta3/ext/iconv/iconv.c#L1424>
------------------------------------------------------------------------
[2018-08-25 12:32:06] cmb@php.net
Oops, wrong bug.
------------------------------------------------------------------------
[2018-08-12 13:17:47] cmb@php.net
> While this seams to be still standard-compliant [â¦]
I'm not sure whether multiple encoded words on the same line
are standard compliant. RFC 2047 states[1]:
| If it is desirable to encode more text than will fit in an
| 'encoded-word' of 75 characters, multiple 'encoded-word's
| (separated by CRLF SPACE) may be used.
> [â¦] but it seems to me that the division by 3 is not necessary
> [â¦]
Simply removing the / 3[2] causes bug48289.phpt[3] to hang.
> A better approach would be to consider all input bytes
> one-by-one and determine the amount of bytes required to encode
> that, and continue adding more characters until the available
> space is filled.
It seems to me that the most appropriate solution would be to
convert the complete input to the desired output charset, and to
apply the encoding afterwards.
[1] <https://tools.ietf.org/html/rfc2047#section-2>
[2] <https://github.com/php/php-src/blob/php-7.3.0beta1/ext/iconv/iconv.c#L1359>
[3] <https://github.com/php/php-src/blob/php-7.3.0beta1/ext/iconv/tests/bug48289.phpt>
------------------------------------------------------------------------
[2017-05-23 08:30:42] php at pointpro dot nl
I went digging in the source code in ext/iconv/iconv.c.
I agree that it results in much longer strings than necessary, the overhead of the charset marker is
high.
The issue is that the number of remaining characters is divided by 3, the maximum amount of bytes
needed to encode any 8-bit character. If the string to-be encoded consists of mainly or solely
ASCII-characters, this results in a encoded word that takes up only one-third of the available
characters.
It is probably a bit more complex than that, but it seems to me that the division by 3 is not
necessary - you could just use the available remaining characters as is: the loop will calculate the
actual number of characters required for the encoded word, and if this exceeds the amount of
remaining character, the amount of remaining characters is reduces to compensate for this. The
downside is that for strings with lots of non-ASCII codes, it will take several more loops to encode
it, so that the performance is decreased slightly.
A better approach would be to consider all input bytes one-by-one and determine the amount of bytes
required to encode that, and continue adding more characters until the available space is filled.
------------------------------------------------------------------------
The remainder of the comments for this report are too long. To view
the rest of the comments, please view the bug report online at
https://bugs.php.net/bug.php?id=66828
--
Edit this bug report at https://bugs.php.net/bug.php?id=66828&edit=1