Bug #66828 [Opn->Ver]: iconv_mime_encode quoted-printable result longer than it should be
| From: | cmb@php.net | Date: | Sun, 12 Aug 2018 13:17:47 +0000 |
| Subject: | Bug #66828 [Opn->Ver]: iconv_mime_encode quoted-printable result longer than it should be | ||
| References: | 1 | Groups: | php.bugs |
| Request: | Send a blank email to php-bugs+get-216732@lists.php.net to get a copy of this message | ||
Edit report at https://bugs.php.net/bug.php?id=66828&edit=1
ID: 66828
Updated by: cmb@php.net
Reported by: st_9876543210 at yahoo dot de
Summary: iconv_mime_encode quoted-printable result longer
than it should be
-Status: Open
+Status: Verified
Type: Bug
Package: ICONV related
PHP Version: Irrelevant
Block user comment: N
Private report: N
New Comment:
> While this seams to be still standard-compliant [â¦]
I'm not sure whether multiple encoded words on the same line
are standard compliant. RFC 2047 states[1]:
| If it is desirable to encode more text than will fit in an
| 'encoded-word' of 75 characters, multiple 'encoded-word's
| (separated by CRLF SPACE) may be used.
> [â¦] but it seems to me that the division by 3 is not necessary
> [â¦]
Simply removing the
/ 3[2] causes bug48289.phpt[3] to hang.
> A better approach would be to consider all input bytes
> one-by-one and determine the amount of bytes required to encode
> that, and continue adding more characters until the available
> space is filled.
It seems to me that the most appropriate solution would be to
convert the complete input to the desired output charset, and to
apply the encoding afterwards.
[1] <https://tools.ietf.org/html/rfc2047#section-2>
[2] <https://github.com/php/php-src/blob/php-7.3.0beta1/ext/iconv/iconv.c#L1359>
[3] <https://github.com/php/php-src/blob/php-7.3.0beta1/ext/iconv/tests/bug48289.phpt>
Previous Comments:
------------------------------------------------------------------------
[2017-05-23 08:30:42] php at pointpro dot nl
I went digging in the source code in ext/iconv/iconv.c.
I agree that it results in much longer strings than necessary, the overhead of the charset marker is
high.
The issue is that the number of remaining characters is divided by 3, the maximum amount of bytes
needed to encode any 8-bit character. If the string to-be encoded consists of mainly or solely
ASCII-characters, this results in a encoded word that takes up only one-third of the available
characters.
It is probably a bit more complex than that, but it seems to me that the division by 3 is not
necessary - you could just use the available remaining characters as is: the loop will calculate the
actual number of characters required for the encoded word, and if this exceeds the amount of
remaining character, the amount of remaining characters is reduces to compensate for this. The
downside is that for strings with lots of non-ASCII codes, it will take several more loops to encode
it, so that the performance is decreased slightly.
A better approach would be to consider all input bytes one-by-one and determine the amount of bytes
required to encode that, and continue adding more characters until the available space is filled.
------------------------------------------------------------------------
[2014-03-05 19:04:39] st_9876543210 at yahoo dot de
Description:
------------
Instead of wrapping a whole quotes-printabled line into =?UTF-8?Q? and ?=, iconv_mime_encode()
sometimes splits a line into 2 parts each surrounded by the mentioned parts. While this seams to be
still standard-compliant it makes the resulting string longer than it should be.
My example would perfectly fit into one line but instead it needs 2 lines.
This also happens when having longer strings (having this issue on every line of the output!) and
regardless of what characters are used.
This could to be caused by the bug fix for https://bugs.php.net/bug.php?id=48289.
Test script:
---------------
<?php
$preferences = array(
"input-charset" => "ISO-8859-1",
"output-charset" => "UTF-8",
"line-length" => 76,
"line-break-chars" => "\n",
"scheme" => "Q"
);
var_dump(iconv_mime_encode("Subject", "Test Test Test Test Test Test Test Test",
$preferences));
Expected result:
----------------
string(67) "Subject: =?UTF-8?Q?Test=20Test=20Test=20Test=20Test=20Test=20Test?="
Actual result:
--------------
string(93) "Subject: =?UTF-8?Q?Test=20Test=20Test=20Tes?==?UTF-8?Q?t=20Test?=
=?UTF-8?Q?=20Test=20Test?="
------------------------------------------------------------------------
--
Edit this bug report at https://bugs.php.net/bug.php?id=66828&edit=1