Re: How should PEAR packages support multiple character sets?
| From: | Jan Schneider | Date: | Thu, 27 Mar 2008 10:45:14 +0000 |
| Subject: | Re: How should PEAR packages support multiple character sets? | ||
| References: | 1 2 3 4 5 | Groups: | php.pear.dev |
| Request: | Send a blank email to pear-dev+get-49563@lists.php.net to get a copy of this message | ||
Zitat von Christian Schmidt <lists@chsc.dk>:
Jan Schneider wrote:The spirit of mbstring.func_overload is broken by design. It is in the same category like register_globals or magic quotes IMHO.Ah, I agree. Packages that offer the luxury of a setCharset() method should continue to do so. I was rather thinking in the line of packages that only support one character set, e.g. . I think these packages should use UTF-8 encoding. Also, I think packages that do offer a setCharset() method should default to UTF-8. As an extra service to people using mbstring function overloading, PEAR packages could honour the mbstring.func_overload setting and assume that strings are encoded using the character set specified by mbstring.internal_encoding, exactly like strlen() etc. do when overloaded. This is similar to what I suggested as Solution 2. I know that this is global-variable-like, but I think it is in the spirit of mbstring.func_overload setting.In your opinion, what kind of package that deals with non-US-ASCII characters should use a character set other than UTF-8?Any package that can safely deal with arbitrary charsets. Since I happen to be the maintainer, take Console_Table for example. It works fine with any charset supported by the mbstring extension, and I don't see a reason why I should limit users to one or two charsets only.
I can see that you have made a great effort to implement multiple character set support in Console_Table. Perhaps your implementations of _strlen() etc. could be moved to a seperate utility package that other packages can use, if they want to offer support for multiple character sets.That would make sense, and these methods in Console_Table are actually taken from Horde's String package which does exactly that: providing charset-safe string manipulation methods.
I can see both approaches having their own merits. I agree it should be one way or the other. OTOH, unless there is a central solution for safely converting charsets in PEAR, each package would have to create its own if they want to accept any user input. A lot of redundancy. Jan. -- Do you need professional PHP or Horde consulting? http://horde.org/consulting/It really depends on the package. Take all the Service_* packages for example that provide interfaces to public APIs. Those usually expect data in a fixed charset. Thus it doesn't make sense for those package to accept data in any charset different to that.I think it makes sense, as long as the user input can be converted to a character set supported by the service. Even if the service only supports e.g. ISO-8859-1, characters not in this character set should be rejected (perhaps by throwing an exception, if the user supplied characters not in this character set) or silently removed (like utf8_decode() does). But if the service supports UTF-7 (like IMAP), UTF-8 or UTF-16, I consider it an implementation detail which encoding is used on the wire, and I don't think this should be exposed to the user. For comparison, the DOM extension always exposes strings as UTF-8, independently of the encoding used in the original XML file.