Req #80689 [Opn]: Add support for incremental encoding conversion
| From: | cmb@php.net | Date: | Sat, 30 Jan 2021 18:26:13 +0000 |
| Subject: | Req #80689 [Opn]: Add support for incremental encoding conversion | ||
| References: | 1 | Groups: | php.bugs |
| Request: | Send a blank email to php-bugs+get-231844@lists.php.net to get a copy of this message | ||
Edit report at https://bugs.php.net/bug.php?id=80689&edit=1
ID: 80689
Updated by: cmb@php.net
Reported by: dhammond at webdevout dot net
Summary: Add support for incremental encoding conversion
Status: Open
Type: Feature/Change Request
Package: mbstring related
PHP Version: 8.0.1
Block user comment: N
Private report: N
New Comment:
Note that iconv conversion stream filters[1] are available.
[1] <https://www.php.net/manual/en/filters.convert.php#filters.convert.iconv>
Previous Comments:
------------------------------------------------------------------------
[2021-01-30 17:50:54] dhammond at webdevout dot net
Description:
------------
Mbstring currently supports converting a complete string of text from one encoding to another, but
it doesn't yet support converting a stream of text incrementally. This is needed in streaming
workflows that process one chunk of bytes at a time.
If you try to use mb_convert_encoding() in a streaming workflow, you run into problems with
multibyte encodings:
1. The chunk might end in the middle of a multibyte sequence, resulting in corruption at the chunk
boundaries.
2. Byte order detection in encodings like UTF-16 gets reset each chunk, meaning it might correctly
interpret the first chunk as UTF-16LE and then incorrectly interpret the next chunk as UTF-16BE.
3. Some special encodings, like BASE64, have unique problems at chunk boundaries. In the case of
BASE64 output encoding, if the input chunk is not a multiple of 3 bytes, then the chunk output will
contain padding characters which should not exist in the middle of base64 data.
These problems would be resolved if we had a way to convert encodings incrementally. The mbstring
module appears to support incremental conversion under the hood, but it doesn't yet expose any
incremental API to userland. Here's an example of how such an API might look:
$context = mb_convert_init('UTF-8', 'UTF-16'); // To convert from UTF-16 to
UTF-8.
while (!$source->feof())
{
$input_chunk = $source->read(8192);
$output_chunk = mb_convert_add($context, $input_chunk, false);
$dest->write($output_chunk);
}
$output_chunk = mb_convert_add($context, '', true);
$dest->write($output_chunk);
In the above example, the third argument of mb_convert_add() is set to true for the final chunk, to
indicate that it should finalize the stream and flush any buffers. In usages that are structured
more like stream filters, it may be more common for this to be called like "$output_chunk =
mb_convert_add($this->context, $input_chunk, $closing);", where the final call may contain
input data that should be added before finalizing the stream.
------------------------------------------------------------------------
--
Edit this bug report at https://bugs.php.net/bug.php?id=80689&edit=1