Re: How should PEAR packages support multiple character sets?

From: Date: Tue, 25 Mar 2008 09:15:07 +0000
Subject: Re: How should PEAR packages support multiple character sets?
References: 1  Groups: php.pear.dev 
Request: Send a blank email to pear-dev+get-49529@lists.php.net to get a copy of this message
Zitat von Christian Schmidt <lists@chsc.dk>:
Strings in PHP5 are simple byte arrays without any information about which character set should be used for decoding the string. When passing a string to a function, the caller is responsible for converting the string to the proper character set. Some built-in PHP functions expects strings to be encoded in ISO-8859-1, while e.g. the DOM extension expects strings to be UTF-8. Some functions are encoding-agnostic, some functions accepts any single-byte encoding and some functions support various encodings based on ini settings. If an application generally assumes that strings are encoded in ISO-8859-1, it is annoying that all strings should be converted using utf8_encode()/utf8_decode() when using the DOM extension. On the other hand, if code assumes that strings are encoded in UTF-8, special care has to be taken when using e.g. strlen(). In general, PEAR packages should just work without assuming too much and without requiring specific ini settings. It would be preferable if all PEAR packages handle character sets the same way. Basically, I see two (realistic) solutions to this challenge: 1. A PEAR package should expect strings to be encoded using UTF-8 2. A PEAR package should expect strings to be encoded using some character set that may be specified by the user at runtime, e.g. PEAR::setCharacterSet('UTF-8'). PEAR classed would use specific to convert user-supplied strings to a particular encoding (e.g. PEAR::convertFromGlobalEncodingToUTF8($string), or to convert internal strings (e.g. fetched from another source over the network) to the globally specified encoding before returning them (e.g. PEAR::convertFromUTF8ToGlobalEncoding($string)). These helper functions should also handle different mbstring ini settings. Solution 1 is easy to explain and implement in PEAR packages, but applications based on ISO-8859-1 must add utf8_encode()/utf8_decode() calls on all strings that are passed to or returned from PEAR methods. Solution 2 is more complicated for PEAR developers but the end user's code may look more elegant. Also, it may be easier to migrate to PHP6 that has native Unicode support. Do you agree that PEAR packages needs a uniform way of dealing with this issue? Do you see other solutions? Which solution do you prefer?
I don't. It really depends on the package. For some packages it simply makes sense to assume utf-8 encoding, especially packages dealing with dom and xml, but probably also others. This has to be documented of course. But it doesn't make sense to force passing a charset to those package that work in a certain context where utf-8 is basically the standard. OTOH, for packages that don't, a global charset setting might still be the wrong approach. PEAR is not a consistent framework, so you can't rely on a global setting. Beside that it has the bad taste of using a global variable. The charset awareness might be available in some packages but not in others. Not being able to precisely distinguish between charsets passes to separate packages is calling for trouble. Thus, the charset should be passed directly to packages, object ctors, or even methods where it fits best for each package. Jan. -- Do you need professional PHP or Horde consulting? http://horde.org/consulting/

« previous php.pear.dev (#49529) next »