Re: Internationalization

From: Date: Fri, 16 Jul 1999 18:19:45 +0000
Subject: Re: Internationalization
References: 1  Groups: php.dev 
Request: Send a blank email to php-dev+get-8605@lists.php.net to get a copy of this message
I would not worry too much about regex stuff. We should be moving to a UTF-handling regex implementation at some point. (Does PCRE do this? I know that the latest Spencer regex stuff that's in Tcl does.) Jim Hironori Sato wrote: > > Hi > (here comes a lengthy mail!) > > I have discussed this issue with Rasmus privately as well as on php-dev > list couple times while ago. > > Thanks to both sgk@happysize.co.jp and tsukada@fminn.nagano.nagano.jp, I18N > version is fairly complete. Since it's becoming much more painfull to > keeping up with main distribution, we are hoping to have the i18n part > incorporated into the main distribution. > > Here are some background on what it does (if you don't care about the nitty > gritty detail, skip this): > ---------------------------------------------------------------------- > (the main concern is to handle Japanese, but many other languages faces > common situation with having multiple character encodings for same language) > > o Problem with current distribution > When dealing with Japanese, you have to accommadate multiple character > codes, mainly SJIS, EUC, JIS, and UTF-8. To make long story short, to > handle Japanese properly, PHP needs to accept any type of encoding and > convert the codes to desired internal encoding. At the sametime, the > output code should be controlled as well. There are few other > modifications needed to accommadate Japanese as well: mail and regex. > > o I18N version > Here are the list of changes in rough cut > - added conversion filters (as we call it) for SJIS, EUC, JIS and UTF-8 > - YY_INPUT in language-scanner is overloaded to include filtering process > - http output is filtered > - mail is modified to comply with RFC for sending email in Japanese > - php3.ini addition to configure how PHP handles various codes > - should be easy to implement other language > - to enable Japanese support, one should compile it with --enable-i18n > > o How common users uses PHP > In php3.ini, a user may set the encoding scheme as such: > > i18n.script_encoding = AUTO > i18n.http_input = AUTO > i18n.internal_encoding = EUC > i18n.http_output = SJIS > > This means, php3 script can be in any encoding, as well as any post or get > data. They will be converted to EUC automagicaly. Now, since the internal > encoding is in EUC, saving any data to MySQL (for an example) will be in > EUC. This is perfectly fine since MySQL will only handle either SJIS or > EUC. Samething applies with any other database, output encoding is > whatever the internal encoding is. At last, output is set to SJIS as > default, but can be controlled via php function. > > o Todo > Regex! This is the most painfull part. Since we allow EUC and SJIS to be > used as internal encoding, regex should be written to accommadate both. > This sucks. The best thing is to find (if any) regex that can handle > unicode, but this means we need to use UTF-8 as internal encoding. This is > where it gets tricky, if internal encoding is UTF, then any output other > than regular http output will be in UTF. We suspect that many applications > that interact with PHP can't handle UTF. The only way to get around this > is to place filtering at each output process (that's alot of work for every > single developers). Any ideas on this? anyone? > Can't compile under Windows. Looking to get more developers in this area. > ---------------------------------------------------------------------- > > We are hoping that by merging with main distribution, we will gain more > Japanese supporters. AFAIK, there is one Japanese book already in publish > that deals with PHP. Core PHP Programming is in process of translation. > So this should be damn good timing. > > This is more than Japanese hack for PHP. We spent long time discussing the > implementation with ultimate goal to be true i18n support. Any multi-byte > user can look at the code and should be able to add their language support. > I think the unicode part should be very interesting to look at. > > I know sgk has CVS access, so he may be able to commit the changes if he > has time. If not I will give it a try. > > Any comments would be greately appreciated! > > Hiro

« previous php.dev (#8605) next »