Re: About IDNA and similar issues... | Call for input
| From: | Matthias Sommerfeld | Date: | Wed, 11 Aug 2004 13:55:14 +0000 |
| Subject: | Re: About IDNA and similar issues... | Call for input | ||
| References: | 1 2 3 4 5 6 | Groups: | php.pear.dev |
| Request: | Send a blank email to pear-dev+get-32549@lists.php.net to get a copy of this message | ||
Dear David,
just a few notes about your post.
Hello Matthias, Unfortunately, I've stopped my work on IDNA for personal reasons. (Sorry..).That's a real pitty. You seemed to be very concerned about the topic.
But, I don't want to miss my comments on extensions. Personally, I think it's good to have both multiple input and output encodings available for unicode representations. This could be done i.e. via a factory() method. Especially mbstring and libiconv are interesting here.I would like to ask the other developers in the list to post their opinion about this here. My personal opinion is, that supporting ther encoding should be done on application level, not hard wired into Net_IDNA, which covers IDNs, not charater conversions.
The overhead is not really big. UCS-4 should, but must not, represent unicode codepoints internally.UCS-4 is part of the Unicode standard. So it can be considered to represent Unicode codepoints.
Full RFC-compliance may need a fewer look on in details. Either you or Markus noted IDNA is just for second-level domains, which is not true. IDNA can occur in every label of a hostname.Don't be concerned. Seems like you didn't test our class. Try this: http://idnaconv.phlymail.de/index.php?decoded=m%C3%BCller.m%C3%BCller.atze.m%C3%BCller&encode=Encode+%3E%3E I am aware, that there's various use cases, where applications need IDNA. The default of the class allows Punycode in every label of a hostname. By using the strict mode, applications can elect to decide on their own, where Punycode is allowed. This behaviour is the most flexible in my eyes.
Maybe someone could take a look on standards compliance here.Maybe this is not an issue one should be worried about. The relevant RFCs are closely implemented. I think we should not tangle poeple again.
I also think Punycode should be split-off as it's own class.This really doesn't make sense. Plus it's conflicting with what you said above. Either you would like to put as much functionality into the class as possible or you wish for modules put together. BTW, RFC3490 states: <cite> IDNA requires that implementations process input strings with Nameprep [NAMEPREP], which is a profile of Stringprep [STRINGPREP], and then with Punycode [PUNYCODE]. Implementations of IDNA MUST fully implement Nameprep and Punycode; neither Nameprep nor Punycode are optional. </cite> Regards Matthias