Re: Internationalization
| From: | Jim Winstead | Date: | Fri, 16 Jul 1999 18:19:45 +0000 |
| Subject: | Re: Internationalization | ||
| References: | 1 | Groups: | php.dev |
| Request: | Send a blank email to php-dev+get-8605@lists.php.net to get a copy of this message | ||
I would not worry too much about regex stuff. We should be moving to a
UTF-handling regex implementation at some point. (Does PCRE do this?
I know that the latest Spencer regex stuff that's in Tcl does.)
Jim
Hironori Sato wrote:
>
> Hi
> (here comes a lengthy mail!)
>
> I have discussed this issue with Rasmus privately as well as on php-dev
> list couple times while ago.
>
> Thanks to both sgk@happysize.co.jp and tsukada@fminn.nagano.nagano.jp, I18N
> version is fairly complete. Since it's becoming much more painfull to
> keeping up with main distribution, we are hoping to have the i18n part
> incorporated into the main distribution.
>
> Here are some background on what it does (if you don't care about the nitty
> gritty detail, skip this):
> ----------------------------------------------------------------------
> (the main concern is to handle Japanese, but many other languages faces
> common situation with having multiple character encodings for same language)
>
> o Problem with current distribution
> When dealing with Japanese, you have to accommadate multiple character
> codes, mainly SJIS, EUC, JIS, and UTF-8. To make long story short, to
> handle Japanese properly, PHP needs to accept any type of encoding and
> convert the codes to desired internal encoding. At the sametime, the
> output code should be controlled as well. There are few other
> modifications needed to accommadate Japanese as well: mail and regex.
>
> o I18N version
> Here are the list of changes in rough cut
> - added conversion filters (as we call it) for SJIS, EUC, JIS and UTF-8
> - YY_INPUT in language-scanner is overloaded to include filtering process
> - http output is filtered
> - mail is modified to comply with RFC for sending email in Japanese
> - php3.ini addition to configure how PHP handles various codes
> - should be easy to implement other language
> - to enable Japanese support, one should compile it with --enable-i18n
>
> o How common users uses PHP
> In php3.ini, a user may set the encoding scheme as such:
>
> i18n.script_encoding = AUTO
> i18n.http_input = AUTO
> i18n.internal_encoding = EUC
> i18n.http_output = SJIS
>
> This means, php3 script can be in any encoding, as well as any post or get
> data. They will be converted to EUC automagicaly. Now, since the internal
> encoding is in EUC, saving any data to MySQL (for an example) will be in
> EUC. This is perfectly fine since MySQL will only handle either SJIS or
> EUC. Samething applies with any other database, output encoding is
> whatever the internal encoding is. At last, output is set to SJIS as
> default, but can be controlled via php function.
>
> o Todo
> Regex! This is the most painfull part. Since we allow EUC and SJIS to be
> used as internal encoding, regex should be written to accommadate both.
> This sucks. The best thing is to find (if any) regex that can handle
> unicode, but this means we need to use UTF-8 as internal encoding. This is
> where it gets tricky, if internal encoding is UTF, then any output other
> than regular http output will be in UTF. We suspect that many applications
> that interact with PHP can't handle UTF. The only way to get around this
> is to place filtering at each output process (that's alot of work for every
> single developers). Any ideas on this? anyone?
> Can't compile under Windows. Looking to get more developers in this area.
> ----------------------------------------------------------------------
>
> We are hoping that by merging with main distribution, we will gain more
> Japanese supporters. AFAIK, there is one Japanese book already in publish
> that deals with PHP. Core PHP Programming is in process of translation.
> So this should be damn good timing.
>
> This is more than Japanese hack for PHP. We spent long time discussing the
> implementation with ultimate goal to be true i18n support. Any multi-byte
> user can look at the code and should be able to add their language support.
> I think the unicode part should be very interesting to look at.
>
> I know sgk has CVS access, so he may be able to commit the changes if he
> has time. If not I will give it a try.
>
> Any comments would be greately appreciated!
>
> Hiro