Bug #78767 [Nab]: mb_convert_case with lower option have error with Turkish capital İ

From: Date: Wed, 13 Nov 2019 15:31:07 +0000
Subject: Bug #78767 [Nab]: mb_convert_case with lower option have error with Turkish capital İ
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-223696@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=78767&edit=1

 ID:                 78767
 Updated by:         requinix@php.net
 Reported by:        meminaydin at cmbilisim dot com
 Summary:            mb_convert_case with lower option have error with
                     Turkish capital İ
 Status:             Not a bug
 Type:               Bug
 Package:            mbstring related
 Operating System:   Debian
 PHP Version:        7.3.11
 Block user comment: N
 Private report:     N

 New Comment:

Please use this bug tracker for comments instead of personal emails.

> You marked it as "not a bug," but it's a bug. Please run the test script for php
> 7.2 and 7.3 at
> http://sandbox.onlinephpfunctions.com/code/a2df81548b6b3bb4b97ea52398fccdae6645c851
> to check the output.

I have explained to you why the change in behavior is not a bug. Please explain to me why you think
it is.


Previous Comments:
------------------------------------------------------------------------
[2019-10-31 21:03:39] requinix@php.net

First, a glossary:
- 0x69 is the Latin 'i' (U+0069 LATIN SMALL LETTER I)
- 0xCC87 is a "combining dot above" (U+0307 COMBINING DOT ABOVE)

Some links:
- https://unicode.org/mail-arch/unicode-ml/Archives-Old/UML009/0619.html
- http://www.i18nguy.com/unicode/turkish-i18n.html
- https://www.nu42.com/2017/02/for-your-eyes-only.html

And some Unicode rules:
- Turkish 'İ' (capital 'I' with dot) is uppercase of Latin 'i'
(lowercase 'i' with dot)
- Latin 'I' (capital 'I' without dot) is uppercase of Turkish 'ı'
(lowercase 'i' without dot)
- Latin 'i' followed by any "above" diacritical mark means that it loses its
normal dot and gains the mark instead

If lowercasing Turkish 'İ' produced Latin 'i' (0x69), then uppercasing it
would produce Latin 'I'. That is incorrect: the dot was lost. By appending the combining
dot, lowercasing produces something that is still visually Latin 'i' but uppercasing can
correctly identify that it should produce the Turkish 'İ' with dot... however it's
not actually possible to know that it should be literally *that* 'İ' (U+0130) because
codepoints don't indicate language, so instead the uppercase is Latin 'I' with
another combining dot.

https://3v4l.org/eWIep

------------------------------------------------------------------------
[2019-10-31 20:27:08] meminaydin at cmbilisim dot com

Description:
------------
mb_convert_case with lower option have a bug. If the string being converted contains turkish capital
i (İ), the output is incorrect.  No problem in versions 7.2 and earlier.

Test script:
---------------
<?php
$str = 'yİy';
echo $tmp = mb_convert_case($str, MB_CASE_LOWER, 'UTF-8'), "\n";

echo implode('', unpack('H*', $tmp)), "\n";

Expected result:
----------------
yiy
796979


Actual result:
--------------
yi̇y
7969cc8779



------------------------------------------------------------------------



--
Edit this bug report at https://bugs.php.net/bug.php?id=78767&edit=1


Thread (5 messages)

« previous php.bugs (#223696) next »