Bug #70475 [Asn]: ext/mbstring/unicode_data.h needs update

From: Date: Tue, 15 Sep 2015 14:59:20 +0000
Subject: Bug #70475 [Asn]: ext/mbstring/unicode_data.h needs update
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-196015@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=70475&edit=1

 ID:                 70475
 Updated by:         laruence@php.net
 Reported by:        cl at exomail dot to
 Summary:            ext/mbstring/unicode_data.h needs update
 Status:             Assigned
 Type:               Bug
 Package:            mbstring related
 Operating System:   all
 PHP Version:        Irrelevant
 Assigned To:        wez
 Block user comment: N
 Private report:     N

 New Comment:

okey, the Unicode_data.h is updated to 8.0.0: https://github.com/php/php-src/commit/e841016df727896342310b579f93dfc55b931caf


Previous Comments:
------------------------------------------------------------------------
[2015-09-14 09:43:52] cl at exomail dot to

The problem has at least two aspects:
1. old data; this can be solved with an update
2. to know what full case folding means

ad 1:
You are not really asking, if php should continue to use the very old/outdated mappings, are you?

ad 2:
If you look at php_unicode.c function case_lookup then you see the php developers had the idea that
one code point is replaced by another code point. 
Hence: You have a problem if you want to replace one code point with more than one codepoint (that
is what has to happen for "FULL CASE FOLDING"). Take a look at 
ftp://ftp.unicode.org/Public/UNIDATA/CaseFolding.txt
to see what Unicode's idea of full case folding is.

00DF; F; 0073 0073; # LATIN SMALL LETTER SHARP S
      ^^ FULL case folding for U+00DF

In short:
case_lookup in php_unicode.c is "defective by design" for doing full case folding. There
are a lot (more than 100) code points that need full case folding including the very common german
"ß".

Even if you decide to give up on full case folding [I really really hope not; how long will php wait
to *fully* support such simple things like strtoupper] at least the update will help for simple case
folding.

------------------------------------------------------------------------
[2015-09-14 03:02:35] laruence@php.net

Hmm, I have tried upgrade the unicode_data.h up to UnicodeData-8.0.0. but seems the behavior is till
the same, thus I am not sure should I do the update..

thanks

------------------------------------------------------------------------
[2015-09-11 13:43:56] cl at exomail dot to

Description:
------------
Looking at github the last update of php-src/ext/mbstring/unicode_data.h was

2010-10-05 42dae97fd49f8d5f5d45c6254794f41fc2b32c88

So the Unicode-Data (for mb_strtoupper, etc.) is FIVE(!) years old.

There is a nice website called http://unicode.org/

I know PHP is playing in a different league than other software languages (which try to follow
unicode changes as closely as possible). But think about looking every 3 or 4 years on this nice
Unicode webpage and include the "newest" changes.

Even for really "old" characters mb_strtoupper does nothing, like for
U+00DF (ß) or U+0149 (ʼn).



Test script:
---------------
$str="\xc3\x9f"; echo $str." upper =>
".mb_strtoupper($str,'UTF-8')."\n";
$str="\xc5\x89"; echo $str." upper =>
".mb_strtoupper($str,'UTF-8')."\n";


Expected result:
----------------
expected output:
ß upper => SS
ʼn upper => ʼN


Actual result:
--------------
output of testscript:
ß upper => ß
ʼn upper => ʼn



------------------------------------------------------------------------



--
Edit this bug report at https://bugs.php.net/bug.php?id=70475&edit=1


Thread (11 messages)

« previous php.bugs (#196015) next »