Re: About IDNA and similar issues...
| From: | David Rech | Date: | Wed, 11 Aug 2004 13:15:11 +0000 |
| Subject: | Re: About IDNA and similar issues... | ||
| References: | 1 2 3 4 5 | Groups: | php.pear.dev |
| Request: | Send a blank email to pear-dev+get-32547@lists.php.net to get a copy of this message | ||
Matthias Sommerfeld wrote:
Dear David, dear Stefan, as mentioned in my last post, I did some changes to Net_IDNA. A summary follows: - Fixed the CS issues discussed in the mailing list - Added support for UCS-4 in string and array represenation as input formats for encode() and output formats for decode(). - Dropped now useless parameter 'use_utf8'. - Added parameter 'encoding' with possible values 'utf8', 'ucs4_string' and 'ucs4_array' for supporting the new formats. This is a global setting. - Added a second parameter to encode() and decode(), where the above values can be set. This parameter affects the current conversion process only. - Fixed a few bugs in the algorithms, which sometimes lead to strange results, especially with Japanese Unciode codepoints. Besides that, we have made the PHP4 version available via CVS. It supports both its original API for maximum backwards compatibility and the Net_IDNA API. Now it's time to talk about possible extensions to the classes for using different encodings or optionally installed PHP IDN extensions. Regards, MatthiasHello Matthias, Unfortunately, I've stopped my work on IDNA for personal reasons. (Sorry..). But, I don't want to miss my comments on extensions. Personally, I think it's good to have both multiple input and output encodings available for unicode representations. This could be done i.e. via a factory() method. Especially mbstring and libiconv are interesting here. The overhead is not really big. UCS-4 should, but must not, represent unicode codepoints internally. Full RFC-compliance may need a fewer look on in details. Either you or Markus noted IDNA is just for second-level domains, which is not true. IDNA can occur in every label of a hostname. Maybe someone could take a look on standards compliance here. Usage of STD3 ASCII rules for hostnames should also be applied. Maybe optionally via a flag. I also think Punycode should be split-off as it's own class. Merging with Stefan's Testcases would also make much sense. -- david rech