Re: First idea for I18N_Punycode
| From: | Stefan Neufeind | Date: | Sun, 09 May 2004 22:19:37 +0000 |
| Subject: | Re: First idea for I18N_Punycode | ||
| References: | 1 | Groups: | php.pear.dev |
| Request: | Send a blank email to pear-dev+get-29050@lists.php.net to get a copy of this message | ||
Hi David,
On 9 May 2004 at 22:40, David Rech wrote:
> Stefan Neufeind wrote:
>
> > as mentioned in a previous email I already had a fully working
> > Punycode-implementation. However, it needed some cleaning up and
> > heavy commenting. Let's see how, together with David, we can get a
> > proposal for that package ready soon.
> >
> > After doing all the commenting, adding extensive tests etc. I tried
> > wrapping up an installable package. To be honest: It took me a while.
> > And currently it's only working with the mbstring-extension (no
> > libiconv yet). Libiconv has problems e.g. converting ISO-8859-1 into
> > UCS-4 ... haven't yet figured out why :-((
> >
> > Anyway, if you want to have a look grab a copy from my current work
> > grab your copy from
> >
> > http://pear.speedpartner.de/packages/I18N_Punycode-0.0.1.tgz
> >
> > On that site you also find autogenerated apidoc (though the API is
> > quite small).
> > The package comes phpdoc-commented and with full unit-tests.
>
> First to note, I ran into some trouble while trying to improve my
> punycode implementation a little less... *gna* I guess it will take a
> while to figure out what went wrong. ;-(
To be honest: When I ran the full RFC-unittests on the initial class
I also discovered some slight mistakes. Unittests definitely rock :-)
> After taking a look on the code you provided, I think we should go on
> with improving and extending your code, because you have most of the
> documentation ready and you're the one who don't need to mess with
> bigger bugs ;-)
Thanks for the flowers.
> A little comment due to correctness:
>
> I saw the ACE-label prefix, xn--, in the I18N_Punycode_ASCII class.
> It's not that important, but the ACE and "put en-/decoding together"
> code is *not* subject of punycode.
Agreed, I was looking at the problem from a "simple" point of view.
But I guess you're right.
> Maybe you already have an extra piece of code for the man IDN stuff? As
> mentioned before, it's not really important, but maybe we should add
> something like I18N_IDN as the package, and I18N_IDN_Punycode as a class
> of the package.
Now that you raised that topic: Maybe I18N_IDN or even I18N_IDNA
(what's the correct term to be used here?) is a better package-name.
However, I don't think that we should over-complicate things: Is a
separate _Punycode-class really necessary? Since its the official,
widely used standard nowadays, do you think we will have
implementations besides Punycode that justify having it flexible in a
separate class? Otherwise, maybe we should save memory and code-
complexity. If you look at the current class you'll notice that the
public methods encode() and decode() do the initial conversions,
split the string into labels, do nameprep (for encoding) and then
call the _punycode-functions. If the class is renamed to I18N_IDN or
I18N_IDNA, don't you think that the current _punycode-functions are
"separate enough"? For sure the punycode-prefix needs to be renamed
to ace-prefix.
> It's not subject of punycode to do label splitting and such stuff, that
> *is* subject of IDNA with the functions referred as idna_to_* e.g. in
> libidn.
Haven't yet dealt myself with the functionality provided by libidn.
Have you? What modifications to the class might be useful?
What I currently see are on the one hand complete implementation in
PHP and on the other using everything from libidn, right? And then we
have the choice of different characterset-conversion-libs if multiple
are available. How could all that needed flexibility best fit into an
easy to use API - especially with the libidn-extension in mind?
> To avoid too big discussion on that, yes I know... most users talking
> about internationalized domains think "uh, idn? isn't that the punycode
> thing?" So we *could* go the way I suggested above but it's no must.
Already commented above.
PS: Maybe for characterset-conversions we should already take a look
at the gnu-recode-functions provided as a addon to php. But haven't
yet evaluated.
David and Johannes, I'd really appreciate your help for finding a
suitable API. Especially with libidn in mind as an option ...
Regards,
Stefan