Bug #70475 [Com]: ext/mbstring/unicode_data.h needs update
| From: | cl at exomail dot to | Date: | Wed, 30 Sep 2015 00:11:16 +0000 |
| Subject: | Bug #70475 [Com]: ext/mbstring/unicode_data.h needs update | ||
| References: | 1 | Groups: | php.bugs |
| Request: | Send a blank email to php-bugs+get-196306@lists.php.net to get a copy of this message | ||
Edit report at https://bugs.php.net/bug.php?id=70475&edit=1
ID: 70475
Comment by: cl at exomail dot to
Reported by: cl at exomail dot to
Summary: ext/mbstring/unicode_data.h needs update
Status: Assigned
Type: Bug
Package: mbstring related
Operating System: all
PHP Version: Irrelevant
Assigned To: wez
Block user comment: N
Private report: N
New Comment:
To summarize:
* The test script expects mbstring to do case *mapping* as defined in
Section 5.18 of http://www.unicode.org/versions/Unicode8.0.0/ch05.pdf
* php-src/ext/mbstring/ucgendat/ucgendat.c generates unicode_data.h
In unicode_data.h the data for 0x00df (Ã) is not there.
Does ucgendat.c use SpecialCasing.txt?
* FAQ 1 in http://unicode.org/faq/casemap_charprop.html:
"Is all of the Unicode case mapping information in UnicodeData.txt?"
"No." Use UnicodeData.txt *and* SpecialCasing.txt!
And Unicode-Standard Section "4.2 Case":
"The single-character mappingsin UnicodeData.txt are insufficient for languages such as
German."
* The data structure
static const unsigned int _uccase_map[] = {
in php's unicode_data.h (IMHO) assumes a one-to-one mapping
/* Starting indexes of the case tables
* UpperIndex = 0
* LowerIndex = _uccase_len[0]
* TitleIndex = LowerIndex + _uccase_len[1] */
Thats the source of the problem.
What other informations are needed to get consensus that this is a bug? [A new bug report with a new
number, really?]
Previous Comments:
------------------------------------------------------------------------
[2015-09-29 22:49:52] cl at exomail dot to
The test script expects mbstring to do case *mapping* as defined in
Section 5.18 "Case Mappings" of "The Unicode Standard"
http://www.unicode.org/versions/Unicode8.0.0/ch05.pdf
There (in the *mapping* section) the "Ã" is even given as example:
toUpperCase("Ã") = "SS"
------------------------------------------------------------------------
[2015-09-29 20:03:35] fsb at thefsb dot org
Data should certainly be updated. But the test script is mistaken.
Case folding is *not* the same as case mapping, see 2nd FAQ here: http://unicode.org/faq/casemap_charprop.html
mbstring provides case mapping, which is what I would expect given the method names and
documentation. The test scripts here, otoh, look like they are expecting case folding.
This might explain why laruence@php.net observed that updating the Unicode data made no difference.
(Btw, from the same FAQ "Beginning with Unicode 5.0, case folding became subject to stability
constraints.")
Unit tests that check if mbstring is using 8.0 could focus instead on things mentioned in the
Unicode 8.0 release announcement, e.g. codepoints becomming assigned. http://blog.unicode.org/2015/06/announcing-unicode-standard-version-80.html
If there's something wrong with mbstring's case conversion, it should be reported in a
separate bug report.
------------------------------------------------------------------------
[2015-09-15 14:59:18] laruence@php.net
okey, the Unicode_data.h is updated to 8.0.0: https://github.com/php/php-src/commit/e841016df727896342310b579f93dfc55b931caf
------------------------------------------------------------------------
[2015-09-14 09:43:52] cl at exomail dot to
The problem has at least two aspects:
1. old data; this can be solved with an update
2. to know what full case folding means
ad 1:
You are not really asking, if php should continue to use the very old/outdated mappings, are you?
ad 2:
If you look at php_unicode.c function case_lookup then you see the php developers had the idea that
one code point is replaced by another code point.
Hence: You have a problem if you want to replace one code point with more than one codepoint (that
is what has to happen for "FULL CASE FOLDING"). Take a look at
ftp://ftp.unicode.org/Public/UNIDATA/CaseFolding.txt
to see what Unicode's idea of full case folding is.
00DF; F; 0073 0073; # LATIN SMALL LETTER SHARP S
^^ FULL case folding for U+00DF
In short:
case_lookup in php_unicode.c is "defective by design" for doing full case folding. There
are a lot (more than 100) code points that need full case folding including the very common german
"Ã".
Even if you decide to give up on full case folding [I really really hope not; how long will php wait
to *fully* support such simple things like strtoupper] at least the update will help for simple case
folding.
------------------------------------------------------------------------
[2015-09-14 03:02:35] laruence@php.net
Hmm, I have tried upgrade the unicode_data.h up to UnicodeData-8.0.0. but seems the behavior is till
the same, thus I am not sure should I do the update..
thanks
------------------------------------------------------------------------
The remainder of the comments for this report are too long. To view
the rest of the comments, please view the bug report online at
https://bugs.php.net/bug.php?id=70475
--
Edit this bug report at https://bugs.php.net/bug.php?id=70475&edit=1