Re: [PEPr] Comment on Networking::Net_IDNA
| From: | Matthias Sommerfeld | Date: | Thu, 22 Jul 2004 12:50:35 +0000 |
| Subject: | Re: [PEPr] Comment on Networking::Net_IDNA | ||
| References: | 1 2 3 | Groups: | php.pear.dev |
| Request: | Send a blank email to pear-dev+get-32248@lists.php.net to get a copy of this message | ||
Johannes Schlueter wrote:
Matthias Sommerfeld wrote:Ok, this is no holy grail for me, too.My idn PECL extension is listed under I18N, too. http://pecl.php.net/idn1) The package name IDNA may be Net related, but doesn't it be more related to Internatiolization at all? I18N_IDN may be a good choice here.I disagree, since IDN(A) is just intended for domain names, not for general internationalisation or localisation purposes. So I think it is really placed best in the Net_ category.
Agreed.Conversion of whole domain names is the thing most people want to do with such a class so imho it's nice to have other possibilities but not2) Punycode Punycode is a single standard. Even if it's only related to IDNA it should (IMHO) go into a single class.Punycode is just one part of encoding / decoding IDNs. The whole story of StringPrep / NamePrep is part of it. So dividing the whole process of an IDN codec into the basic parts of Punycode conversion, Nameprep and so on doesn't make sense to me.
required. The idn extension currently has only two functions, either convert a Unicode-string holding a domain name to IDN-Ascii or the other way round. Somwhen I'll go and offer the TLD-checking with the extension (some TLDs allow only limited parts of the Unicode set) which is offered by the GNU libidn I'm using and if someone asks I'll export the other functions for stringprep or Punycode but I don't see them with high priorityThat's something I price similarly. Nice to have, not required. And of course, again a huge data source like the NamePrep data.
One big issue of PHP in the moment is, that it is not Unicode enabled. As long as the textual data you process is in the range of ASCII and ISO-Latin-1 everything goes well. But having to deal with textual data from other charsets or even pure Unicode data you are on your own. So using parse_url() on a "real world" domain name in Unicode will probably fail.imho splitting the a URL into hostname, URI, etc. can easily be done with parse_url() so this has not to be part of the class.3) Input encodings I'd like to see support for many more input encodings than just the unicode-compatible ones. Especially "Multibyte string" (mbstring) ind "libiconv" (iconv) extension are interesting here. Also "recode" may need a quick look on.I disagree. This should be part of the application using Net_IDNA. I also tend to remove the routines of snipping the input string into the domain name parts, protocol and query string again, since this approach seems to be to much overhead in the sense of IDN. IDN is specified for second level domains only, so extracting the second level domain from a complete link in the face of the fact, that there's various Unicode codepoints representing a dot is ugly and prone to fail. But this last comment is also open for discussion :)
Yeah, that would be real nice, since I considered my first steps to create a PHP written IDN class as a fallback solution for those, who cannot use libidn or the like (how many C extensions for PHP are there already? 3 or 4? I guess each and every on of them has their own API). So I would really love to adjust Net_IDNA's API in a way, that any application handling IDN can switch easily between ext_idn and PEAR_idn.Besides the things pointed out above I consider the class being quite stable and complete.Maybe we could look over it to create some similar API for ext_idn and PEAR_idn so a user can easily select wether he can use a C written extension or if he has to use the PHP version but don't need to change too much of his code.
johannes-- Matthias Sommerfeld phlyLabs mailto:mso@phlylabs.de http://phlylabs.de