Re: First idea for I18N_Punycode

From: Date: Mon, 10 May 2004 08:31:06 +0000
Subject: Re: First idea for I18N_Punycode
References: 1 2  Groups: php.pear.dev 
Request: Send a blank email to pear-dev+get-29055@lists.php.net to get a copy of this message
Zitat von Stefan Neufeind <stefan@neufeind.net>:
Maybe you already have an extra piece of code for the man IDN stuff? As mentioned before, it's not really important, but maybe we should add something like I18N_IDN as the package, and I18N_IDN_Punycode as a class of the package.
Now that you raised that topic: Maybe I18N_IDN or even I18N_IDNA (what's the correct term to be used here?) is a better package-name. However, I don't think that we should over-complicate things: Is a separate _Punycode-class really necessary? Since its the official, widely used standard nowadays, do you think we will have implementations besides Punycode that justify having it flexible in a separate class? Otherwise, maybe we should save memory and code-
Why not keep forward compatibility? It's not much work, and if there will be any other encoding than punycode in the future, you'd be prepared.
complexity. If you look at the current class you'll notice that the public methods encode() and decode() do the initial conversions, split the string into labels, do nameprep (for encoding) and then call the _punycode-functions. If the class is renamed to I18N_IDN or I18N_IDNA, don't you think that the current _punycode-functions are "separate enough"? For sure the punycode-prefix needs to be renamed to ace-prefix.
I think two separate classes make sense.
It's not subject of punycode to do label splitting and such stuff, that *is* subject of IDNA with the functions referred as idna_to_* e.g. in libidn.
Haven't yet dealt myself with the functionality provided by libidn. Have you? What modifications to the class might be useful? What I currently see are on the one hand complete implementation in PHP and on the other using everything from libidn, right? And then we have the choice of different characterset-conversion-libs if multiple are available. How could all that needed flexibility best fit into an easy to use API - especially with the libidn-extension in mind?
The idn ext currently only expose two methods, to encode and to decode strings. I'd extend the en-/decode() methods of your package to allow an optional third parameter that defines the encoding methods to use and default it to punycode. If punycode is specified as the encoding method (explicitely or by the default value), check if the idn ext is available and use its methods then. Otherwise use the userland implementation.
To avoid too big discussion on that, yes I know... most users talking about internationalized domains think "uh, idn? isn't that the punycode thing?" So we *could* go the way I suggested above but it's no must.
Already commented above. PS: Maybe for characterset-conversions we should already take a look at the gnu-recode-functions provided as a addon to php. But haven't yet evaluated.
From my experience iconv and mbstring are the only reliable extensions to use for charset conversion. Recode is not suitable because it takes other charset names than the other extensions and transliterates by default (correct me someone if I'm wrong.). Jan. -- Do you need professional PHP or Horde consulting? http://horde.org/consulting/

« previous php.pear.dev (#29055) next »