Bug #66828 [Com]: iconv_mime_encode quoted-printable result longer than it should be

From: Date: Tue, 23 May 2017 08:30:49 +0000
Subject: Bug #66828 [Com]: iconv_mime_encode quoted-printable result longer than it should be
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-209232@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=66828&edit=1 ID: 66828 Comment by: php at pointpro dot nl Reported by: st_9876543210 at yahoo dot de Summary: iconv_mime_encode quoted-printable result longer than it should be Status: Open Type: Bug Package: ICONV related PHP Version: Irrelevant Block user comment: N Private report: N New Comment: I went digging in the source code in ext/iconv/iconv.c. I agree that it results in much longer strings than necessary, the overhead of the charset marker is high. The issue is that the number of remaining characters is divided by 3, the maximum amount of bytes needed to encode any 8-bit character. If the string to-be encoded consists of mainly or solely ASCII-characters, this results in a encoded word that takes up only one-third of the available characters. It is probably a bit more complex than that, but it seems to me that the division by 3 is not necessary - you could just use the available remaining characters as is: the loop will calculate the actual number of characters required for the encoded word, and if this exceeds the amount of remaining character, the amount of remaining characters is reduces to compensate for this. The downside is that for strings with lots of non-ASCII codes, it will take several more loops to encode it, so that the performance is decreased slightly. A better approach would be to consider all input bytes one-by-one and determine the amount of bytes required to encode that, and continue adding more characters until the available space is filled. Previous Comments: ------------------------------------------------------------------------ [2014-03-05 19:04:39] st_9876543210 at yahoo dot de Description: ------------ Instead of wrapping a whole quotes-printabled line into =?UTF-8?Q? and ?=, iconv_mime_encode() sometimes splits a line into 2 parts each surrounded by the mentioned parts. While this seams to be still standard-compliant it makes the resulting string longer than it should be. My example would perfectly fit into one line but instead it needs 2 lines. This also happens when having longer strings (having this issue on every line of the output!) and regardless of what characters are used. This could to be caused by the bug fix for https://bugs.php.net/bug.php?id=48289. Test script: --------------- <?php $preferences = array( "input-charset" => "ISO-8859-1", "output-charset" => "UTF-8", "line-length" => 76, "line-break-chars" => "\n", "scheme" => "Q" ); var_dump(iconv_mime_encode("Subject", "Test Test Test Test Test Test Test Test", $preferences)); Expected result: ---------------- string(67) "Subject: =?UTF-8?Q?Test=20Test=20Test=20Test=20Test=20Test=20Test?=" Actual result: -------------- string(93) "Subject: =?UTF-8?Q?Test=20Test=20Test=20Tes?==?UTF-8?Q?t=20Test?= =?UTF-8?Q?=20Test=20Test?=" ------------------------------------------------------------------------ -- Edit this bug report at https://bugs.php.net/bug.php?id=66828&edit=1

« previous php.bugs (#209232) next »