Edit report at https://bugs.php.net/bug.php?id=70475&edit=1
ID: 70475
Updated by: nikic@php.net
Reported by: cl at exomail dot to
Summary: ext/mbstring/unicode_data.h needs update
-Status: Assigned
+Status: Closed
Type: Bug
Package: mbstring related
Operating System: all
PHP Version: Irrelevant
Assigned To: wez
Block user comment: N
Private report: N
New Comment:
Closing here as unicode data has been updated and bug #70609 deals with the full case mapping.
I've also updated the data to Unicode 10.0 in PHP 7.2.
Previous Comments:
------------------------------------------------------------------------
[2015-09-30 18:02:26] fsb at thefsb dot org
(Thanks for the new bug report, cl at exomail dot to.)
Before closing this one, I think it might be wise for mbstring to use Unicode 7 in PHP 7.0 rather
than Unicode 8.
First: PCRE is stuck on Unicode 7 and I think it's going to stay that way. PCRE2 is on Unicode
8 but there's no sign of PHP adopting it.
Second: intl ext is using ICU 55.1 which is Unicode 7. ICU 56-rc is available but it seems likely
7.0 will go to GA with 55.1.
So it might make sense for PHP 7.0 to be consistent and use Unicode 7 across the board.
------------------------------------------------------------------------
[2015-09-30 14:01:04] cl at exomail dot to
https://bugs.php.net/bug.php?id=70609
------------------------------------------------------------------------
[2015-09-30 00:54:08] fsb at thefsb dot org
That's a great summary of the mapping problem, cl at exomail dot to.
The title of this bug is "ext/mbstring/unicode_data.h needs update" and laruence@php.net
has since updated it to UCD 8.0, so I think single-to-multi mapping is a separate bug report /
feature request.
------------------------------------------------------------------------
[2015-09-30 00:11:12] cl at exomail dot to
To summarize:
* The test script expects mbstring to do case *mapping* as defined in
Section 5.18 of http://www.unicode.org/versions/Unicode8.0.0/ch05.pdf
* php-src/ext/mbstring/ucgendat/ucgendat.c generates unicode_data.h
In unicode_data.h the data for 0x00df (Ã) is not there.
Does ucgendat.c use SpecialCasing.txt?
* FAQ 1 in http://unicode.org/faq/casemap_charprop.html:
"Is all of the Unicode case mapping information in UnicodeData.txt?"
"No." Use UnicodeData.txt *and* SpecialCasing.txt!
And Unicode-Standard Section "4.2 Case":
"The single-character mappingsin UnicodeData.txt are insufficient for languages such as
German."
* The data structure
static const unsigned int _uccase_map[] = {
in php's unicode_data.h (IMHO) assumes a one-to-one mapping
/* Starting indexes of the case tables
* UpperIndex = 0
* LowerIndex = _uccase_len[0]
* TitleIndex = LowerIndex + _uccase_len[1] */
Thats the source of the problem.
What other informations are needed to get consensus that this is a bug? [A new bug report with a new
number, really?]
------------------------------------------------------------------------
[2015-09-29 22:49:52] cl at exomail dot to
The test script expects mbstring to do case *mapping* as defined in
Section 5.18 "Case Mappings" of "The Unicode Standard"
http://www.unicode.org/versions/Unicode8.0.0/ch05.pdf
There (in the *mapping* section) the "Ã" is even given as example:
toUpperCase("Ã") = "SS"
------------------------------------------------------------------------
The remainder of the comments for this report are too long. To view
the rest of the comments, please view the bug report online at
https://bugs.php.net/bug.php?id=70475
--
Edit this bug report at https://bugs.php.net/bug.php?id=70475&edit=1