Re: multi-byte awareness for formatted_print.c

From: Date: Wed, 10 Oct 2001 14:14:59 +0000
Subject: Re: multi-byte awareness for formatted_print.c
References: 1 2 3  Groups: php.dev 
Request: Send a blank email to php-dev+get-67691@lists.php.net to get a copy of this message
At 16:16 10-10-01, Wez Furlong wrote:
On 10/10/01, "Zeev Suraski" <zeev@zend.com> wrote: <warning> I wasn't following this thread, just an out-of-context comment </warning> OK :-) That shouldn't be a problem. str++ advances the address according to the type of its base... But a multi-byte character can be 1 or more bytes, depending on the character and the encoding, so you can't just say each character is 2 bytes long unless you are sure that the encoding says that each character is always 2 bytes long. It's great isn't it? In my experience with Japanese encodings you get into all kinds of problems if you are not careful. To be able to use str++ in that way, we would need to convert to string into an appropriate wide character format before working with it, and then convert it back afterwards.
Oh, well, that greatly depends on which kind of encoding we end up working with. Some encodings have fixed width, others do not... Zeev

« previous php.dev (#67691) next »