note 105755 deleted from function.mb-strtolower by danbrown

From: Date: Tue, 13 Sep 2011 00:42:01 +0000
Subject: note 105755 deleted from function.mb-strtolower by danbrown
References: 1  Groups: php.notes 
Request: Send a blank email to php-notes+get-183127@lists.php.net to get a copy of this message
Note Submitter: akniep at rayo dot info ---- In addition to my last comment: Since I was not finding any proper solution in the internet on how to map all UTF8-strings to their lowercase counterpart in PHP, I offer the following hard-coded extended mb_strtolower function for UTF-8 strings: The function wraps the existing function mb_strtolower() and additionally replaces uppercase UTF8-characters for which there is a lowercase representation. Since there is no proper Unicode uppercase and lowercase character-table in the internet that I was able to find, I checked the first million UTF8-characters against the Google-search and -KeywordTool and identified the following 49 characters as uppercase-characters, not being replaced by mb_strtolower, but having a UTF8 lowercase counterpart. <?php // the numbers in the in-line-comments display the characters' Unicode code-points (CP). function strtolower_utf8_extended( $utf8_string ) { $additional_replacements = array ( "Dž" => "dž" // CP 453 -> 454 , "Lj" => "lj" // CP 456 -> 457 , "Nj" => "nj" // CP 459 -> 460 , "Dz" => "dz" // CP 498 -> 499 , "Ϸ" => "ϸ" // CP 1015 -> 1016 , "Ϲ" => "ϲ" // CP 1017 -> 1010 , "Ϻ" => "ϻ" // CP 1018 -> 1019 , "ᾋ" => "ᾃ" // CP 8075 -> 8067 , "ᾌ" => "ᾄ" // CP 8076 -> 8068 , "ᾍ" => "ᾅ" // CP 8077 -> 8069 , "ᾎ" => "ᾆ" // CP 8078 -> 8070 , "ᾏ" => "ᾇ" // CP 8079 -> 8071 , "ᾘ" => "ᾐ" // CP 8088 -> 8080 , "ᾙ" => "ᾑ" // CP 8089 -> 8081 , "ᾚ" => "ᾒ" // CP 8090 -> 8082 , "ᾛ" => "ᾓ" // CP 8091 -> 8083 , "ᾜ" => "ᾔ" // CP 8092 -> 8084 , "ᾝ" => "ᾕ" // CP 8093 -> 8085 , "ᾞ" => "ᾖ" // CP 8094 -> 8086 , "ᾟ" => "ᾗ" // CP 8095 -> 8087 , "ᾨ" => "ᾠ" // CP 8104 -> 8096 , "ᾩ" => "ᾡ" // CP 8105 -> 8097 , "ᾪ" => "ᾢ" // CP 8106 -> 8098 , "ᾫ" => "ᾣ" // CP 8107 -> 8099 , "ᾬ" => "ᾤ" // CP 8108 -> 8100 , "ᾭ" => "ᾥ" // CP 8109 -> 8101 , "ᾮ" => "ᾦ" // CP 8110 -> 8102 , "ᾯ" => "ᾧ" // CP 8111 -> 8103 , "ᾼ" => "ᾳ" // CP 8124 -> 8115 , "ῌ" => "ῃ" // CP 8140 -> 8131 , "ῼ" => "ῳ" // CP 8188 -> 8179 , "Ⅰ" => "ⅰ" // CP 8544 -> 8560 , "Ⅱ" => "ⅱ" // CP 8545 -> 8561 , "Ⅲ" => "ⅲ" // CP 8546 -> 8562 , "Ⅳ" => "ⅳ" // CP 8547 -> 8563 , "Ⅴ" => "ⅴ" // CP 8548 -> 8564 , "Ⅵ" => "ⅵ" // CP 8549 -> 8565 , "Ⅶ" => "ⅶ" // CP 8550 -> 8566 , "Ⅷ" => "ⅷ" // CP 8551 -> 8567 , "Ⅸ" => "ⅸ" // CP 8552 -> 8568 , "Ⅹ" => "ⅹ" // CP 8553 -> 8569 , "Ⅺ" => "ⅺ" // CP 8554 -> 8570 , "Ⅻ" => "ⅻ" // CP 8555 -> 8571 , "Ⅼ" => "ⅼ" // CP 8556 -> 8572 , "Ⅽ" => "ⅽ" // CP 8557 -> 8573 , "Ⅾ" => "ⅾ" // CP 8558 -> 8574 , "Ⅿ" => "ⅿ" // CP 8559 -> 8575 , "𐐦" => "𐑎" // CP 66598 -> 66638 , "𐐧" => "𐑏" // CP 66599 -> 66639 ); $utf8_string = mb_strtolower( $utf8_string, "UTF-8"); $utf8_string = strtr( $utf8_string, $additional_replacements ); return $utf8_string; } //strtolower_utf8_extended() ?>

« previous php.notes (#183127) next »