Re: [RFC] IntlCharsetDetector
| From: | Tom Worster | Date: | Wed, 27 Apr 2016 14:34:04 +0000 |
| Subject: | Re: [RFC] IntlCharsetDetector | ||
| References: | 1 2 3 | Groups: | php.internals |
| Request: | Send a blank email to internals+get-92839@lists.php.net to get a copy of this message | ||
On 4/26/16 12:10 PM, Sara Golemon wrote:
On Tue, Apr 26, 2016 at 2:06 AM, Yasuo Ohgaki <yohgaki@ohgaki.net> wrote:Why do you expect that? When I researched this problem some years ago I had the impression a number of attempted solutions had been published and abandoned. I took this to mean that there was a learning experience that ended with the understanding that it's insoluble. That's why I'm curious if you know of ongoing efforts in ICU. I took a look and saw little activity in the last 10 years.Things might have been changed, but as you've mentioned encoding detection is unstable and ICU is poor compared to mbstring's detection at least for Japanese encodings.For me, the difference is that I expect further work to be done on improving ICU,
while I lack that confidence for mbstring. If the API is in place early on, the library can improve underneath it to the point it becomes more trustworthy later, but still be usable on older versions of PHP (linked against newer libicu).How would it becomes more trustworthy? A way to make it trustworthy would need to exist. And somebody would have to work on it. Tom