Edit report at https://bugs.php.net/bug.php?id=70475&edit=1
ID: 70475
Comment by: cl at exomail dot to
Reported by: cl at exomail dot to
Summary: ext/mbstring/unicode_data.h needs update
Status: Assigned
Type: Bug
Package: mbstring related
Operating System: all
PHP Version: Irrelevant
Assigned To: wez
Block user comment: N
Private report: N
New Comment:
The problem has at least two aspects:
1. old data; this can be solved with an update
2. to know what full case folding means
ad 1:
You are not really asking, if php should continue to use the very old/outdated mappings, are you?
ad 2:
If you look at php_unicode.c function case_lookup then you see the php developers had the idea that
one code point is replaced by another code point.
Hence: You have a problem if you want to replace one code point with more than one codepoint (that
is what has to happen for "FULL CASE FOLDING"). Take a look at
ftp://ftp.unicode.org/Public/UNIDATA/CaseFolding.txt
to see what Unicode's idea of full case folding is.
00DF; F; 0073 0073; # LATIN SMALL LETTER SHARP S
^^ FULL case folding for U+00DF
In short:
case_lookup in php_unicode.c is "defective by design" for doing full case folding. There
are a lot (more than 100) code points that need full case folding including the very common german
"Ã".
Even if you decide to give up on full case folding [I really really hope not; how long will php wait
to *fully* support such simple things like strtoupper] at least the update will help for simple case
folding.
Previous Comments:
------------------------------------------------------------------------
[2015-09-14 03:02:35] laruence@php.net
Hmm, I have tried upgrade the unicode_data.h up to UnicodeData-8.0.0. but seems the behavior is till
the same, thus I am not sure should I do the update..
thanks
------------------------------------------------------------------------
[2015-09-11 13:43:56] cl at exomail dot to
Description:
------------
Looking at github the last update of php-src/ext/mbstring/unicode_data.h was
2010-10-05 42dae97fd49f8d5f5d45c6254794f41fc2b32c88
So the Unicode-Data (for mb_strtoupper, etc.) is FIVE(!) years old.
There is a nice website called http://unicode.org/
I know PHP is playing in a different league than other software languages (which try to follow
unicode changes as closely as possible). But think about looking every 3 or 4 years on this nice
Unicode webpage and include the "newest" changes.
Even for really "old" characters mb_strtoupper does nothing, like for
U+00DF (Ã) or U+0149 (Å).
Test script:
---------------
$str="\xc3\x9f"; echo $str." upper =>
".mb_strtoupper($str,'UTF-8')."\n";
$str="\xc5\x89"; echo $str." upper =>
".mb_strtoupper($str,'UTF-8')."\n";
Expected result:
----------------
expected output:
à upper => SS
Šupper => ʼN
Actual result:
--------------
output of testscript:
à upper => Ã
Å upper => Å
------------------------------------------------------------------------
--
Edit this bug report at https://bugs.php.net/bug.php?id=70475&edit=1