Re: Multiple char encodings for site internationalzation
| From: | Hironori Sato | Date: | Tue, 03 Aug 1999 07:58:44 +0000 |
| Subject: | Re: Multiple char encodings for site internationalzation | ||
| References: | 1 | Groups: | php.dev |
| Request: | Send a blank email to php-dev+get-9472@lists.php.net to get a copy of this message | ||
Here goes I18N issue again.
At 16:34 99/7/29 -0700, Alvin Engler wrote:
> Is there any move towards implementing support for more non-latin
> character sets in the works? (the iso 8859-X series and windows Code
> pages)
>
> I see the utf8_encode and decode as well as the convert_cyr_encoding
> functions, but what about middle eastern and asian languages such as
> arabic etc...
The Japanese developers has implemented some modification onto PHP. To see
the modification, check out the patch:
ftp://ftp.happysize.co.jp/php-ja-jp/php-3.0.7jp-beta2.patch1.tar.gz
Don't have English document to go along with it, but patch should be self
explanatory. 3.0.12 version is prepared, but needs some testing.
> I am in charge of the technical aspects of translating a site into
> arabic, and need to serve users who are only capable of displaying one
> of 3 arabic encodings (windows code page 1256, iso 8859-6 or UTF-8) as
> well as those who do not have an arabic browser at all (dynamic gifs
> created with tt fonts).
>
> with such a function I would simply store my base data in UTF-8 form and
> convert it using the function based on the user selection or
> automatically based on thier user-agent.
>
> I don't feel proficient enough (my c skills are pretty weak) to write my
> own module, but it seems that with such a library already in existance
> it should be a simple task to write such a function....
So, Arabic has three encodings? We got four!!! The above patch can take
any PHP scripts and HTTP post/get, then converted to one fixed encoding
('internal encoding'). Also, you can send http output in any of four
encoding. If you want to implement Arabic, you can make your own
convertors like: filt_1256_utf8, filt_1256_8859, filt_8859_utf8, etc...
Just like we have filt_jis_eucjp, filt_eucjp_jis, filt_jis_utf8, and so on.
Can you handle it? If so, we can definitely work togather.
The problem with building PHP function to convert between different code is
the trouble that script writer has to go through dealing with encoding
conversion manually. With above implementation, there is no need to worry
about the encoding. Unless you want to fiddle with encoding, you can
simply set the handling of various output/input via php3.ini. That's all.
I really feel i18n is becoming very important issue with PHP. What do
everyone at PHP-DEV think? Since there weren't much response from the last
email, I'm pretty disappointed... :(
This brings me to a question. I read the licensing term on zend stuff, but
couldn't come to a solid conclusion. If someone modifies part of Zend, can
it be implemented without causing any problem with licensing? It's very
minor part of language-scanner. Is it a big hassel? > Zend team.
Thanks in advance.
Hiro
//--- Hironori Sato --------------------------------------- KB9HAD ---
// satoh@jpnnet.com http://staff.jpnnet.com/satoh/
// Japanese Network http://www.jpnnet.com/