Re: Multiple char encodings for site internationalzation

From: Date: Thu, 05 Aug 1999 02:27:38 +0000
Subject: Re: Multiple char encodings for site internationalzation
References: 1 2 3  Groups: php.dev 
Request: Send a blank email to php-dev+get-9540@lists.php.net to get a copy of this message
Japanese, Arabic, Russians... Who else has problem dealing with multiple charsets? I'm guessing Chinese (both simplified and traditional) has the same problem too. Though haven't heard anyone complain about it. At 11:43 99/8/4 +0400, Vadim Kolontsov wrote: > In php3 cyrillic has same problems with UTF-8: > > utf8_encode doesn't know anything about russian (cyrillic) charsets >(actually you can't use any charset except ISO-8859-1 in php3's >utf8_encode()); >and convert_cyr_string() doesn't know anything about UTF-8 :) That sounds very usefull... :-) > I've made a little patch (see atatchment) which allow you to use cyrillic >charsets in utf8_encode()/utf8_decode(). It supports five cyrillic charsets: >iso-8859-5, koi8-r, cp-866, windows-1251, x-mac-cyrillic; it depends on >cyr_convert.c (because I didn't want to mainain two copies of >encoding tables - in xml.c and cyr_convert.c). Is it easy to detect above 5 sets automatically? If so, you can implement the conversion with the way Japanese PHP does as long as you can write the charset conversion routine. I will try to whip out a document on how to do this in near future. I will get back with this when I get the documentation ready. Hmmm, instead of throwing large utf-8 table into PHP distribution, would it be nicer to have some sort of 'language module' for each language? If one needs Japanese support, use 'Japanese module', people in need of Russian can grab 'Russian module', and so on. Just an idea. Hiro //--- Hironori Sato --------------------------------------- KB9HAD --- // satoh@jpnnet.com http://staff.jpnnet.com/satoh/ // Japanese Network http://www.jpnnet.com/

« previous php.dev (#9540) next »