note 105753 modified in function.mb-strtolower by danbrown
| From: | danbrown@php.net | Date: | Tue, 13 Sep 2011 14:25:17 +0000 |
| Subject: | note 105753 modified in function.mb-strtolower by danbrown | ||
| References: | 1 | Groups: | php.notes |
| Request: | Send a blank email to php-notes+get-183142@lists.php.net to get a copy of this message | ||
Please, note that when using with UTF-8 mb_strtolower will only convert upper case characters to
lower case which are marked with the Unicode property "Upper case letter"
("Lu"). However, there are also letters such as "Letter numbers" (Unicode
property "Nl") that also have lower case and upper case variants. These characters will
not be converted be mb_strtolower!
Example:
The Roman letters â
, â
¡, â
¢, ..., â
¯ (UTF-8 code points 8544 through 8559) also
exist in their respective lower case variants â
°, â
±, â
², ..., â
¿ (UTF-8 code points
8560 through 8575) and should, in my opinion, also be converted by mb_strtolower, but they are not!
Big internet-companies (like Google) do match both variants as semantically equal (since the
representations only differ in case).
Since I was not finding any proper solution in the internet on how to map all UTF8-strings to their
lowercase counterpart in PHP, I offer the following hard-coded extended mb_strtolower function for
UTF-8 strings:
The function wraps the existing function mb_strtolower() and additionally replaces uppercase
UTF8-characters for which there is a lowercase representation. Since there is no proper Unicode
uppercase and lowercase character-table in the internet that I was able to find, I checked the first
million UTF8-characters against the Google-search and -KeywordTool and identified the following 78
characters as uppercase-characters, not being replaced by mb_strtolower, but having a UTF8 lowercase
counterpart.
<?php
//the numbers in the in-line-comments display the characters' Unicode code-points (CP).
function strtolower_utf8_extended( $utf8_string )
{
$additional_replacements = array
( "Ç
" => "Ç" // 453 -> 454
, "Ç" => "Ç" // 456 -> 457
, "Ç" => "Ç" // 459 -> 460
, "Dz" => "dz" // 498 -> 499
, "Ϸ" => "ϸ" // 1015 -> 1016
, "Ϲ" => "ϲ" // 1017 -> 1010
, "Ϻ" => "ϻ" // 1018 -> 1019
, "á¾" => "á¾" // 8072 -> 8064
, "á¾" => "á¾" // 8073 -> 8065
, "á¾" => "á¾" // 8074 -> 8066
, "á¾" => "á¾" // 8075 -> 8067
, "á¾" => "á¾" // 8076 -> 8068
, "á¾" => "á¾
" // 8077 -> 8069
, "á¾" => "á¾" // 8078 -> 8070
, "á¾" => "á¾" // 8079 -> 8071
, "á¾" => "á¾" // 8088 -> 8080
, "á¾" => "á¾" // 8089 -> 8081
, "á¾" => "á¾" // 8090 -> 8082
, "á¾" => "á¾" // 8091 -> 8083
, "á¾" => "á¾" // 8092 -> 8084
, "á¾" => "á¾" // 8093 -> 8085
, "á¾" => "á¾" // 8094 -> 8086
, "á¾" => "á¾" // 8095 -> 8087
, "ᾨ" => "ᾠ" // 8104 -> 8096
, "ᾩ" => "ᾡ" // 8105 -> 8097
, "ᾪ" => "ᾢ" // 8106 -> 8098
, "ᾫ" => "ᾣ" // 8107 -> 8099
, "ᾬ" => "ᾤ" // 8108 -> 8100
, "á¾" => "á¾¥" // 8109 -> 8101
, "ᾮ" => "ᾦ" // 8110 -> 8102
, "ᾯ" => "ᾧ" // 8111 -> 8103
, "á¾¼" => "á¾³" // 8124 -> 8115
, "á¿" => "á¿" // 8140 -> 8131
, "ῼ" => "ῳ" // 8188 -> 8179
, "â
" => "â
°" // 8544 -> 8560
, "â
¡" => "â
±" // 8545 -> 8561
, "â
¢" => "â
²" // 8546 -> 8562
, "â
£" => "â
³" // 8547 -> 8563
, "â
¤" => "â
´" // 8548 -> 8564
, "â
¥" => "â
µ" // 8549 -> 8565
, "â
¦" => "â
¶" // 8550 -> 8566
, "â
§" => "â
·" // 8551 -> 8567
, "â
¨" => "â
¸" // 8552 -> 8568
, "â
©" => "â
¹" // 8553 -> 8569
, "â
ª" => "â
º" // 8554 -> 8570
, "â
«" => "â
»" // 8555 -> 8571
, "â
¬" => "â
¼" // 8556 -> 8572
, "â
" => "â
½" // 8557 -> 8573
, "â
®" => "â
¾" // 8558 -> 8574
, "â
¯" => "â
¿" // 8559 -> 8575
, "â¶" => "â" // 9398 -> 9424
, "â·" => "â" // 9399 -> 9425
, "â¸" => "â" // 9400 -> 9426
, "â¹" => "â" // 9401 -> 9427
, "âº" => "â" // 9402 -> 9428
, "â»" => "â" // 9403 -> 9429
, "â¼" => "â" // 9404 -> 9430
, "â½" => "â" // 9405 -> 9431
, "â¾" => "â" // 9406 -> 9432
, "â¿" => "â" // 9407 -> 9433
, "â" => "â" // 9408 -> 9434
, "â" => "â" // 9409 -> 9435
, "â" => "â" // 9410 -> 9436
, "â" => "â" // 9411 -> 9437
, "â" => "â" // 9412 -> 9438
, "â
" => "â" // 9413 -> 9439
, "â" => "â " // 9414 -> 9440
, "â" => "â¡" // 9415 -> 9441
, "â" => "â¢" // 9416 -> 9442
, "â" => "â£" // 9417 -> 9443
, "â" => "â¤" // 9418 -> 9444
, "â" => "â¥" // 9419 -> 9445
, "â" => "â¦" // 9420 -> 9446
, "â" => "â§" // 9421 -> 9447
, "â" => "â¨" // 9422 -> 9448
, "â" => "â©" // 9423 -> 9449
, "ð¦" => "ð" // 66598 -> 66638
, "ð§" => "ð" // 66599 -> 66639
);
$utf8_string = mb_strtolower( $utf8_string, "UTF-8");
$utf8_string = strtr( $utf8_string, $additional_replacements );
return $utf8_string;
} //strtolower_utf8_extended()
?>
--was--
Please, note that when using with UTF-8 mb_strtolower will only convert upper case characters to
lower case which are marked with the Unicode property "Upper case letter"
("Lu"). However, there are also letters such as "Letter numbers" (Unicode
property "Nl") that also have lower case and upper case variants. These characters will
not be converted be mb_strtolower!
Example:
The Roman letters â
, â
¡, â
¢, ..., â
¯ (UTF-8 code points 8544 through 8559) also
exist in their respective lower case variants â
°, â
±, â
², ..., â
¿ (UTF-8 code points
8560 through 8575) and should, in my opinion, also be converted by mb_strtolower, but they are not!
Big internet-companies (like Google) do match both variants as semantically equal (since the
representations only differ in case).
http://php.net/manual/en/function.mb-strtolower.php