How should PEAR packages support multiple character sets?
| From: | Christian Schmidt | Date: | Mon, 24 Mar 2008 23:45:22 +0000 |
| Subject: | How should PEAR packages support multiple character sets? | ||
| Groups: | php.pear.dev | ||
| Request: | Send a blank email to pear-dev+get-49528@lists.php.net to get a copy of this message | ||
Strings in PHP5 are simple byte arrays without any information about which character set should be used for decoding the string. When passing a string to a function, the caller is responsible for converting the string to the proper character set.
Some built-in PHP functions expects strings to be encoded in ISO-8859-1, while e.g. the DOM extension expects strings to be UTF-8. Some functions are encoding-agnostic, some functions accepts any single-byte encoding and some functions support various encodings based on ini settings.
If an application generally assumes that strings are encoded in ISO-8859-1, it is annoying that all strings should be converted using utf8_encode()/utf8_decode() when using the DOM extension. On the other hand, if code assumes that strings are encoded in UTF-8, special care has to be taken when using e.g. strlen().
In general, PEAR packages should just work without assuming too much and without requiring specific ini settings. It would be preferable if all PEAR packages handle character sets the same way.
Basically, I see two (realistic) solutions to this challenge:
1. A PEAR package should expect strings to be encoded using UTF-8
2. A PEAR package should expect strings to be encoded using some character set that may be specified by the user at runtime, e.g. PEAR::setCharacterSet('UTF-8'). PEAR classed would use specific to convert user-supplied strings to a particular encoding (e.g. PEAR::convertFromGlobalEncodingToUTF8($string), or to convert internal strings (e.g. fetched from another source over the network) to the globally specified encoding before returning them (e.g. PEAR::convertFromUTF8ToGlobalEncoding($string)). These helper functions should also handle different mbstring ini settings.
Solution 1 is easy to explain and implement in PEAR packages, but applications based on ISO-8859-1 must add utf8_encode()/utf8_decode() calls on all strings that are passed to or returned from PEAR methods.
Solution 2 is more complicated for PEAR developers but the end user's code may look more elegant. Also, it may be easier to migrate to PHP6 that has native Unicode support.
Do you agree that PEAR packages needs a uniform way of dealing with this issue? Do you see other solutions? Which solution do you prefer?
Christian