Unicode and MBCS support using ICU/xIUA
| From: | Carl W. Brown | Date: | Fri, 12 Oct 2001 23:31:08 +0000 |
| Subject: | Unicode and MBCS support using ICU/xIUA | ||
| Groups: | php.dev | ||
| Request: | Send a blank email to php-dev+get-67873@lists.php.net to get a copy of this message | ||
About a year and a half ago I developed code to enable PHP3 to support
Unicode using ICU. These was some initial interest but it died fast so I
did not look into what it would take to develop a PHP4 solution. The
original problem was adding UTF-16 (UCS-2) support to PHP. In the mean time
I have perfected the ICU interface code and have recently made it available
is open source code.
It is thread safe cross-platform and not only supports all forms of Unicode
(UTF-8, UTF-16 & UTF-32) but it supports code page data with the same code.
It is great for browsers because it can use the same functions to process
code page and Unicode data dynamically so that the code does not change if
you are using UTF-8 for one browser and EUC-JP for another. It has new
functions to make PHP charset handling easier. No there is no need to add
16 bit data types to PHP.
I added the Unicode support to PHP3 but making it a semi-resident module. It
also required some minor changes to the HTTP header processing. With the
new xIUA code the changes to PHP are far less. It also make PHP thread safe
in that each thread can have different locales. This is something that
setlocale does not provide.
PHP internal data types are interanlly char *. This would be how I would
imagine implementing Unicode in PHP. The internal data is stored as char *
strings even if it is UTF-16 or UTF-32 data. xIUA will sort it out
automatically. If the data in a string is converted that the date type
attribute changes to match to internal format but the string is still char
*.
This would involve a minimal set of changes and would be backwards
compatible. If for example I collate a UTF-8 string with a UTF-16 string the
UTF-8 string is fast transformed to a UTF-16 string in a stack storage
mechnisim and it is collated. If the UTF-16 string is concatenated to UTF-8
data it is converted to UTF-8. The user does not care it is just data. If
send a EUC-JP script page to a Shift_JIS browser it gets converted.
xIUA contains support for most of the poplar browser code pages so it can be
used for code page support as well. It also contains functions that will
convert strftime date time formats to ICU formats using the ICU internal
data files to produce locale correct formats. It ahs reotines to parse
accpt-language strings or fint the best characters set for a specific locale
basied on accept-charset strings. It has Appahes mads that allow you to
organiage PHP scripts that have been localeized into different
subdirectories by locale and it will optionaly override the mod_mine
processing. It has code that will foce Apache to call the PHP per directory
processing so that setting can vary by directory.
I am not sure how much of this has changed in PHP4 but the code is there if
you need it.
Carl