Re: How should PEAR packages support multiple character sets?

From: Date: Thu, 27 Mar 2008 10:45:14 +0000
Subject: Re: How should PEAR packages support multiple character sets?
References: 1 2 3 4 5  Groups: php.pear.dev 
Request: Send a blank email to pear-dev+get-49563@lists.php.net to get a copy of this message
Zitat von Christian Schmidt <lists@chsc.dk>:
Jan Schneider wrote:
In your opinion, what kind of package that deals with non-US-ASCII characters should use a character set other than UTF-8?
Any package that can safely deal with arbitrary charsets. Since I happen to be the maintainer, take Console_Table for example. It works fine with any charset supported by the mbstring extension, and I don't see a reason why I should limit users to one or two charsets only.
Ah, I agree. Packages that offer the luxury of a setCharset() method should continue to do so. I was rather thinking in the line of packages that only support one character set, e.g. . I think these packages should use UTF-8 encoding. Also, I think packages that do offer a setCharset() method should default to UTF-8. As an extra service to people using mbstring function overloading, PEAR packages could honour the mbstring.func_overload setting and assume that strings are encoded using the character set specified by mbstring.internal_encoding, exactly like strlen() etc. do when overloaded. This is similar to what I suggested as Solution 2. I know that this is global-variable-like, but I think it is in the spirit of mbstring.func_overload setting.
The spirit of mbstring.func_overload is broken by design. It is in the same category like register_globals or magic quotes IMHO.
I can see that you have made a great effort to implement multiple character set support in Console_Table. Perhaps your implementations of _strlen() etc. could be moved to a seperate utility package that other packages can use, if they want to offer support for multiple character sets.
That would make sense, and these methods in Console_Table are actually taken from Horde's String package which does exactly that: providing charset-safe string manipulation methods.
It really depends on the package. Take all the Service_* packages for example that provide interfaces to public APIs. Those usually expect data in a fixed charset. Thus it doesn't make sense for those package to accept data in any charset different to that.
I think it makes sense, as long as the user input can be converted to a character set supported by the service. Even if the service only supports e.g. ISO-8859-1, characters not in this character set should be rejected (perhaps by throwing an exception, if the user supplied characters not in this character set) or silently removed (like utf8_decode() does). But if the service supports UTF-7 (like IMAP), UTF-8 or UTF-16, I consider it an implementation detail which encoding is used on the wire, and I don't think this should be exposed to the user. For comparison, the DOM extension always exposes strings as UTF-8, independently of the encoding used in the original XML file.
I can see both approaches having their own merits. I agree it should be one way or the other. OTOH, unless there is a central solution for safely converting charsets in PEAR, each package would have to create its own if they want to accept any user input. A lot of redundancy. Jan. -- Do you need professional PHP or Horde consulting? http://horde.org/consulting/

« previous php.pear.dev (#49563) next »