Bug #70475 [Com]: ext/mbstring/unicode_data.h needs update
| From: | fsb at thefsb dot org | Date: | Tue, 29 Sep 2015 20:03:40 +0000 |
| Subject: | Bug #70475 [Com]: ext/mbstring/unicode_data.h needs update | ||
| References: | 1 | Groups: | php.bugs |
| Request: | Send a blank email to php-bugs+get-196303@lists.php.net to get a copy of this message | ||
Edit report at https://bugs.php.net/bug.php?id=70475&edit=1
ID: 70475
Comment by: fsb at thefsb dot org
Reported by: cl at exomail dot to
Summary: ext/mbstring/unicode_data.h needs update
Status: Assigned
Type: Bug
Package: mbstring related
Operating System: all
PHP Version: Irrelevant
Assigned To: wez
Block user comment: N
Private report: N
New Comment:
Data should certainly be updated. But the test script is mistaken.
Case folding is *not* the same as case mapping, see 2nd FAQ here: http://unicode.org/faq/casemap_charprop.html
mbstring provides case mapping, which is what I would expect given the method names and
documentation. The test scripts here, otoh, look like they are expecting case folding.
This might explain why laruence@php.net observed that updating the Unicode data made no difference.
(Btw, from the same FAQ "Beginning with Unicode 5.0, case folding became subject to stability
constraints.")
Unit tests that check if mbstring is using 8.0 could focus instead on things mentioned in the
Unicode 8.0 release announcement, e.g. codepoints becomming assigned. http://blog.unicode.org/2015/06/announcing-unicode-standard-version-80.html
If there's something wrong with mbstring's case conversion, it should be reported in a
separate bug report.
Previous Comments:
------------------------------------------------------------------------
[2015-09-15 14:59:18] laruence@php.net
okey, the Unicode_data.h is updated to 8.0.0: https://github.com/php/php-src/commit/e841016df727896342310b579f93dfc55b931caf
------------------------------------------------------------------------
[2015-09-14 09:43:52] cl at exomail dot to
The problem has at least two aspects:
1. old data; this can be solved with an update
2. to know what full case folding means
ad 1:
You are not really asking, if php should continue to use the very old/outdated mappings, are you?
ad 2:
If you look at php_unicode.c function case_lookup then you see the php developers had the idea that
one code point is replaced by another code point.
Hence: You have a problem if you want to replace one code point with more than one codepoint (that
is what has to happen for "FULL CASE FOLDING"). Take a look at
ftp://ftp.unicode.org/Public/UNIDATA/CaseFolding.txt
to see what Unicode's idea of full case folding is.
00DF; F; 0073 0073; # LATIN SMALL LETTER SHARP S
^^ FULL case folding for U+00DF
In short:
case_lookup in php_unicode.c is "defective by design" for doing full case folding. There
are a lot (more than 100) code points that need full case folding including the very common german
"Ã".
Even if you decide to give up on full case folding [I really really hope not; how long will php wait
to *fully* support such simple things like strtoupper] at least the update will help for simple case
folding.
------------------------------------------------------------------------
[2015-09-14 03:02:35] laruence@php.net
Hmm, I have tried upgrade the unicode_data.h up to UnicodeData-8.0.0. but seems the behavior is till
the same, thus I am not sure should I do the update..
thanks
------------------------------------------------------------------------
[2015-09-11 13:43:56] cl at exomail dot to
Description:
------------
Looking at github the last update of php-src/ext/mbstring/unicode_data.h was
2010-10-05 42dae97fd49f8d5f5d45c6254794f41fc2b32c88
So the Unicode-Data (for mb_strtoupper, etc.) is FIVE(!) years old.
There is a nice website called http://unicode.org/
I know PHP is playing in a different league than other software languages (which try to follow
unicode changes as closely as possible). But think about looking every 3 or 4 years on this nice
Unicode webpage and include the "newest" changes.
Even for really "old" characters mb_strtoupper does nothing, like for
U+00DF (Ã) or U+0149 (Å).
Test script:
---------------
$str="\xc3\x9f"; echo $str." upper =>
".mb_strtoupper($str,'UTF-8')."\n";
$str="\xc5\x89"; echo $str." upper =>
".mb_strtoupper($str,'UTF-8')."\n";
Expected result:
----------------
expected output:
à upper => SS
Šupper => ʼN
Actual result:
--------------
output of testscript:
à upper => Ã
Å upper => Å
------------------------------------------------------------------------
--
Edit this bug report at https://bugs.php.net/bug.php?id=70475&edit=1