Req #80689 [Com]: Add support for incremental encoding conversion

From: Date: Mon, 08 Feb 2021 07:50:31 +0000
Subject: Req #80689 [Com]: Add support for incremental encoding conversion
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-232004@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=80689&edit=1

 ID:                 80689
 Comment by:         marshawright19791 at gmail dot com
 Reported by:        dhammond at webdevout dot net
 Summary:            Add support for incremental encoding conversion
 Status:             Open
 Type:               Feature/Change Request
 Package:            mbstring related
 PHP Version:        8.0.1
 Block user comment: N
 Private report:     N

 New Comment:

Classes that don't support this are serialized and stored with the: Each data row contains a
name, and value. The row also contains a: https://www.alaskasworld.me/ type or mimetype. Type
corresponds to a .NET class that support: text/value conversion through the TypeConverter
architecture. Classes that don't support this are serialized and stored .


Previous Comments:
------------------------------------------------------------------------
[2021-01-30 18:26:12] cmb@php.net

Note that iconv conversion stream filters[1] are available.

[1] <https://www.php.net/manual/en/filters.convert.php#filters.convert.iconv>

------------------------------------------------------------------------
[2021-01-30 17:50:54] dhammond at webdevout dot net

Description:
------------
Mbstring currently supports converting a complete string of text from one encoding to another, but
it doesn't yet support converting a stream of text incrementally. This is needed in streaming
workflows that process one chunk of bytes at a time.

If you try to use mb_convert_encoding() in a streaming workflow, you run into problems with
multibyte encodings:

1. The chunk might end in the middle of a multibyte sequence, resulting in corruption at the chunk
boundaries.
2. Byte order detection in encodings like UTF-16 gets reset each chunk, meaning it might correctly
interpret the first chunk as UTF-16LE and then incorrectly interpret the next chunk as UTF-16BE.
3. Some special encodings, like BASE64, have unique problems at chunk boundaries. In the case of
BASE64 output encoding, if the input chunk is not a multiple of 3 bytes, then the chunk output will
contain padding characters which should not exist in the middle of base64 data.

These problems would be resolved if we had a way to convert encodings incrementally. The mbstring
module appears to support incremental conversion under the hood, but it doesn't yet expose any
incremental API to userland. Here's an example of how such an API might look:

$context = mb_convert_init('UTF-8', 'UTF-16'); // To convert from UTF-16 to
UTF-8.

while (!$source->feof())
{
  $input_chunk = $source->read(8192);
  $output_chunk = mb_convert_add($context, $input_chunk, false);
  $dest->write($output_chunk);
}

$output_chunk = mb_convert_add($context, '', true);
$dest->write($output_chunk);

In the above example, the third argument of mb_convert_add() is set to true for the final chunk, to
indicate that it should finalize the stream and flush any buffers. In usages that are structured
more like stream filters, it may be more common for this to be called like "$output_chunk =
mb_convert_add($this->context, $input_chunk, $closing);", where the final call may contain
input data that should be added before finalizing the stream.



------------------------------------------------------------------------



--
Edit this bug report at https://bugs.php.net/bug.php?id=80689&edit=1


Thread (3 messages)

« previous php.bugs (#232004) next »