Re: multi-byte awareness for formatted_print.c

From: Date: Wed, 10 Oct 2001 22:27:14 +0000
Subject: Re: multi-byte awareness for formatted_print.c
References: 1 2  Groups: php.dev 
Request: Send a blank email to php-dev+get-67725@lists.php.net to get a copy of this message
Wez Furlong wrote: > > On 10/10/01, "Zeev Suraski" <zeev@zend.com> wrote: > > <warning> > > I wasn't following this thread, just an out-of-context comment > > </warning> > OK :-) > > > That shouldn't be a problem. str++ advances the address according to the > > type of its base... > > But a multi-byte character can be 1 or more bytes, depending on the character > and the encoding, so you can't just say each character is 2 bytes long unless > you are sure that the encoding says that each character is always 2 bytes long. No, a multi-byte _character_ is a fixed number of bits, its _octet encoding_ may be 1 or more bytes. :-) > It's great isn't it? > > In my experience with Japanese encodings you get into all kinds of problems > if you are not careful. You mean such as always remember to put an extra space after a Shift-JIS string in a form value field? :-) > To be able to use str++ in that way, we would need to convert to string > into an appropriate wide character format before working with it, and > then convert it back afterwards. What's str++, some C++ thingie? - Stig

« previous php.dev (#67725) next »