Re: [PEPr] Comment on Networking::Net_IDNA

From: Date: Thu, 22 Jul 2004 12:50:35 +0000
Subject: Re: [PEPr] Comment on Networking::Net_IDNA
References: 1 2 3  Groups: php.pear.dev 
Request: Send a blank email to pear-dev+get-32248@lists.php.net to get a copy of this message
Johannes Schlueter wrote:
Matthias Sommerfeld wrote:
1) The package name IDNA may be Net related, but doesn't it be more related to Internatiolization at all? I18N_IDN may be a good choice here.
I disagree, since IDN(A) is just intended for domain names, not for general internationalisation or localisation purposes. So I think it is really placed best in the Net_ category.
My idn PECL extension is listed under I18N, too. http://pecl.php.net/idn
Ok, this is no holy grail for me, too.
2) Punycode Punycode is a single standard. Even if it's only related to IDNA it should (IMHO) go into a single class.
Punycode is just one part of encoding / decoding IDNs. The whole story of StringPrep / NamePrep is part of it. So dividing the whole process of an IDN codec into the basic parts of Punycode conversion, Nameprep and so on doesn't make sense to me.
Conversion of whole domain names is the thing most people want to do with such a class so imho it's nice to have other possibilities but not
Agreed.
required. The idn extension currently has only two functions, either convert a Unicode-string holding a domain name to IDN-Ascii or the other way round. Somwhen I'll go and offer the TLD-checking with the extension (some TLDs allow only limited parts of the Unicode set) which is offered by the GNU libidn I'm using and if someone asks I'll export the other functions for stringprep or Punycode but I don't see them with high priority
That's something I price similarly. Nice to have, not required. And of course, again a huge data source like the NamePrep data.
3) Input encodings I'd like to see support for many more input encodings than just the unicode-compatible ones. Especially "Multibyte string" (mbstring) ind "libiconv" (iconv) extension are interesting here. Also "recode" may need a quick look on.
I disagree. This should be part of the application using Net_IDNA. I also tend to remove the routines of snipping the input string into the domain name parts, protocol and query string again, since this approach seems to be to much overhead in the sense of IDN. IDN is specified for second level domains only, so extracting the second level domain from a complete link in the face of the fact, that there's various Unicode codepoints representing a dot is ugly and prone to fail. But this last comment is also open for discussion :)
imho splitting the a URL into hostname, URI, etc. can easily be done with parse_url() so this has not to be part of the class.
One big issue of PHP in the moment is, that it is not Unicode enabled. As long as the textual data you process is in the range of ASCII and ISO-Latin-1 everything goes well. But having to deal with textual data from other charsets or even pure Unicode data you are on your own. So using parse_url() on a "real world" domain name in Unicode will probably fail.
Besides the things pointed out above I consider the class being quite stable and complete.
Maybe we could look over it to create some similar API for ext_idn and PEAR_idn so a user can easily select wether he can use a C written extension or if he has to use the PHP version but don't need to change too much of his code.
Yeah, that would be real nice, since I considered my first steps to create a PHP written IDN class as a fallback solution for those, who cannot use libidn or the like (how many C extensions for PHP are there already? 3 or 4? I guess each and every on of them has their own API). So I would really love to adjust Net_IDNA's API in a way, that any application handling IDN can switch easily between ext_idn and PEAR_idn.
johannes
-- Matthias Sommerfeld phlyLabs mailto:mso@phlylabs.de http://phlylabs.de

« previous php.pear.dev (#32248) next »