Req #80689 [Com]: Add support for incremental encoding conversion
Edit report at https://bugs.php.net/bug.php?id=80689&edit=1
ID: 80689
Comment by: marshawright19791 at gmail dot com
Reported by: dhammond at webdevout dot net
Summary: Add support for incremental encoding conversion
Status: Open
Type: Feature/Change Request
Package: mbstring related
PHP Version: 8.0.1
Block user comment: N
Private report: N
New Comment:
Classes that don't support this are serialized and stored with the: Each data row contains a
name, and value. The row also contains a: https://www.alaskasworld.me/ type or mimetype. Type
corresponds to a .NET class that support: text/value conversion through the TypeConverter
architecture. Classes that don't support this are serialized and stored .
Previous Comments:
------------------------------------------------------------------------
[2021-01-30 18:26:12] cmb@php.net
Note that iconv conversion stream filters[1] are available.
[1] <https://www.php.net/manual/en/filters.convert.php#filters.convert.iconv>
------------------------------------------------------------------------
[2021-01-30 17:50:54] dhammond at webdevout dot net
Description:
------------
Mbstring currently supports converting a complete string of text from one encoding to another, but
it doesn't yet support converting a stream of text incrementally. This is needed in streaming
workflows that process one chunk of bytes at a time.
If you try to use mb_convert_encoding() in a streaming workflow, you run into problems with
multibyte encodings:
1. The chunk might end in the middle of a multibyte sequence, resulting in corruption at the chunk
boundaries.
2. Byte order detection in encodings like UTF-16 gets reset each chunk, meaning it might correctly
interpret the first chunk as UTF-16LE and then incorrectly interpret the next chunk as UTF-16BE.
3. Some special encodings, like BASE64, have unique problems at chunk boundaries. In the case of
BASE64 output encoding, if the input chunk is not a multiple of 3 bytes, then the chunk output will
contain padding characters which should not exist in the middle of base64 data.
These problems would be resolved if we had a way to convert encodings incrementally. The mbstring
module appears to support incremental conversion under the hood, but it doesn't yet expose any
incremental API to userland. Here's an example of how such an API might look:
$context = mb_convert_init('UTF-8', 'UTF-16'); // To convert from UTF-16 to
UTF-8.
while (!$source->feof())
{
$input_chunk = $source->read(8192);
$output_chunk = mb_convert_add($context, $input_chunk, false);
$dest->write($output_chunk);
}
$output_chunk = mb_convert_add($context, '', true);
$dest->write($output_chunk);
In the above example, the third argument of mb_convert_add() is set to true for the final chunk, to
indicate that it should finalize the stream and flush any buffers. In usages that are structured
more like stream filters, it may be more common for this to be called like "$output_chunk =
mb_convert_add($this->context, $input_chunk, $closing);", where the final call may contain
input data that should be added before finalizing the stream.
------------------------------------------------------------------------
--
Edit this bug report at https://bugs.php.net/bug.php?id=80689&edit=1
Thread (3 messages)