Bug #70475 [Com]: ext/mbstring/unicode_data.h needs update

From: Date: Wed, 30 Sep 2015 00:11:16 +0000
Subject: Bug #70475 [Com]: ext/mbstring/unicode_data.h needs update
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-196306@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=70475&edit=1 ID: 70475 Comment by: cl at exomail dot to Reported by: cl at exomail dot to Summary: ext/mbstring/unicode_data.h needs update Status: Assigned Type: Bug Package: mbstring related Operating System: all PHP Version: Irrelevant Assigned To: wez Block user comment: N Private report: N New Comment: To summarize: * The test script expects mbstring to do case *mapping* as defined in Section 5.18 of http://www.unicode.org/versions/Unicode8.0.0/ch05.pdf * php-src/ext/mbstring/ucgendat/ucgendat.c generates unicode_data.h In unicode_data.h the data for 0x00df (ß) is not there. Does ucgendat.c use SpecialCasing.txt? * FAQ 1 in http://unicode.org/faq/casemap_charprop.html: "Is all of the Unicode case mapping information in UnicodeData.txt?" "No." Use UnicodeData.txt *and* SpecialCasing.txt! And Unicode-Standard Section "4.2 Case": "The single-character mappingsin UnicodeData.txt are insufficient for languages such as German." * The data structure static const unsigned int _uccase_map[] = { in php's unicode_data.h (IMHO) assumes a one-to-one mapping /* Starting indexes of the case tables * UpperIndex = 0 * LowerIndex = _uccase_len[0] * TitleIndex = LowerIndex + _uccase_len[1] */ Thats the source of the problem. What other informations are needed to get consensus that this is a bug? [A new bug report with a new number, really?] Previous Comments: ------------------------------------------------------------------------ [2015-09-29 22:49:52] cl at exomail dot to The test script expects mbstring to do case *mapping* as defined in Section 5.18 "Case Mappings" of "The Unicode Standard" http://www.unicode.org/versions/Unicode8.0.0/ch05.pdf There (in the *mapping* section) the "ß" is even given as example: toUpperCase("ß") = "SS" ------------------------------------------------------------------------ [2015-09-29 20:03:35] fsb at thefsb dot org Data should certainly be updated. But the test script is mistaken. Case folding is *not* the same as case mapping, see 2nd FAQ here: http://unicode.org/faq/casemap_charprop.html mbstring provides case mapping, which is what I would expect given the method names and documentation. The test scripts here, otoh, look like they are expecting case folding. This might explain why laruence@php.net observed that updating the Unicode data made no difference. (Btw, from the same FAQ "Beginning with Unicode 5.0, case folding became subject to stability constraints.") Unit tests that check if mbstring is using 8.0 could focus instead on things mentioned in the Unicode 8.0 release announcement, e.g. codepoints becomming assigned. http://blog.unicode.org/2015/06/announcing-unicode-standard-version-80.html If there's something wrong with mbstring's case conversion, it should be reported in a separate bug report. ------------------------------------------------------------------------ [2015-09-15 14:59:18] laruence@php.net okey, the Unicode_data.h is updated to 8.0.0: https://github.com/php/php-src/commit/e841016df727896342310b579f93dfc55b931caf ------------------------------------------------------------------------ [2015-09-14 09:43:52] cl at exomail dot to The problem has at least two aspects: 1. old data; this can be solved with an update 2. to know what full case folding means ad 1: You are not really asking, if php should continue to use the very old/outdated mappings, are you? ad 2: If you look at php_unicode.c function case_lookup then you see the php developers had the idea that one code point is replaced by another code point. Hence: You have a problem if you want to replace one code point with more than one codepoint (that is what has to happen for "FULL CASE FOLDING"). Take a look at ftp://ftp.unicode.org/Public/UNIDATA/CaseFolding.txt to see what Unicode's idea of full case folding is. 00DF; F; 0073 0073; # LATIN SMALL LETTER SHARP S ^^ FULL case folding for U+00DF In short: case_lookup in php_unicode.c is "defective by design" for doing full case folding. There are a lot (more than 100) code points that need full case folding including the very common german "ß". Even if you decide to give up on full case folding [I really really hope not; how long will php wait to *fully* support such simple things like strtoupper] at least the update will help for simple case folding. ------------------------------------------------------------------------ [2015-09-14 03:02:35] laruence@php.net Hmm, I have tried upgrade the unicode_data.h up to UnicodeData-8.0.0. but seems the behavior is till the same, thus I am not sure should I do the update.. thanks ------------------------------------------------------------------------ The remainder of the comments for this report are too long. To view the rest of the comments, please view the bug report online at https://bugs.php.net/bug.php?id=70475 -- Edit this bug report at https://bugs.php.net/bug.php?id=70475&edit=1

« previous php.bugs (#196306) next »