Re: Multiple char encodings for site internationalzation
| From: | Hironori Sato | Date: | Thu, 05 Aug 1999 02:27:38 +0000 |
| Subject: | Re: Multiple char encodings for site internationalzation | ||
| References: | 1 2 3 | Groups: | php.dev |
| Request: | Send a blank email to php-dev+get-9540@lists.php.net to get a copy of this message | ||
Japanese, Arabic, Russians... Who else has problem dealing with multiple
charsets? I'm guessing Chinese (both simplified and traditional) has the
same problem too. Though haven't heard anyone complain about it.
At 11:43 99/8/4 +0400, Vadim Kolontsov wrote:
> In php3 cyrillic has same problems with UTF-8:
>
> utf8_encode doesn't know anything about russian (cyrillic) charsets
>(actually you can't use any charset except ISO-8859-1 in php3's
>utf8_encode());
>and convert_cyr_string() doesn't know anything about UTF-8 :)
That sounds very usefull... :-)
> I've made a little patch (see atatchment) which allow you to use cyrillic
>charsets in utf8_encode()/utf8_decode(). It supports five cyrillic charsets:
>iso-8859-5, koi8-r, cp-866, windows-1251, x-mac-cyrillic; it depends on
>cyr_convert.c (because I didn't want to mainain two copies of
>encoding tables - in xml.c and cyr_convert.c).
Is it easy to detect above 5 sets automatically? If so, you can implement
the conversion with the way Japanese PHP does as long as you can write the
charset conversion routine. I will try to whip out a document on how to do
this in near future. I will get back with this when I get the
documentation ready.
Hmmm, instead of throwing large utf-8 table into PHP distribution, would it
be nicer to have some sort of 'language module' for each language? If one
needs Japanese support, use 'Japanese module', people in need of Russian
can grab 'Russian module', and so on. Just an idea.
Hiro
//--- Hironori Sato --------------------------------------- KB9HAD ---
// satoh@jpnnet.com http://staff.jpnnet.com/satoh/
// Japanese Network http://www.jpnnet.com/