Re: multi-byte awareness for formatted_print.c
| From: | Stig S. Bakken | Date: | Wed, 10 Oct 2001 22:27:14 +0000 |
| Subject: | Re: multi-byte awareness for formatted_print.c | ||
| References: | 1 2 | Groups: | php.dev |
| Request: | Send a blank email to php-dev+get-67725@lists.php.net to get a copy of this message | ||
Wez Furlong wrote:
>
> On 10/10/01, "Zeev Suraski" <zeev@zend.com> wrote:
> > <warning>
> > I wasn't following this thread, just an out-of-context comment
> > </warning>
> OK :-)
>
> > That shouldn't be a problem. str++ advances the address according to the
> > type of its base...
>
> But a multi-byte character can be 1 or more bytes, depending on the character
> and the encoding, so you can't just say each character is 2 bytes long unless
> you are sure that the encoding says that each character is always 2 bytes long.
No, a multi-byte _character_ is a fixed number of bits, its _octet
encoding_ may be 1 or more bytes. :-)
> It's great isn't it?
>
> In my experience with Japanese encodings you get into all kinds of problems
> if you are not careful.
You mean such as always remember to put an extra space after a Shift-JIS
string in a form value field? :-)
> To be able to use str++ in that way, we would need to convert to string
> into an appropriate wide character format before working with it, and
> then convert it back afterwards.
What's str++, some C++ thingie?
- Stig