note 111815 deleted from function.utf8-decode by cmb
| From: | cmb@php.net | Date: | Sun, 11 Dec 2016 11:31:52 +0000 |
| Subject: | note 111815 deleted from function.utf8-decode by cmb | ||
| References: | 1 | Groups: | php.notes |
| Request: | Send a blank email to php-notes+get-208529@lists.php.net to get a copy of this message | ||
Note Submitter: Nitrogen
----
Hi all. I have written this comprehensive UTF-8 decoder as I needed a way to split any UTF-8 string
into nice and useable pieces of information, such as an array of code points that make up the input
string regardless if it's UTF-8 or not.
Here's what's good about this function:
1) It splits a UTF-8 encoded string into an array of code points.
2) It supports from 1 byte (7 bit) to 6 byte (31 bit) UTF-8 characters.
3) It checks the validity of the UTF-8 input string and provides an array of erroneous UTF-8
offsets.
4) Provides you with the character length of the UTF-8 string.
<?php
function get_utf_8($input, $include_raw = FALSE) {
$result = array(
'length'=>0,
'errorbytes'=>array(),
'codepoints'=>array()
);
$codepoints = array();
$len = strlen($input);
for($i=0;$i<$len;$i++) {
$c = ord($input{$i});
$msb = ($c>>7)&1;
$bytes = 1;
if(!$msb)
$codepoints[]= $include_raw?array($c, chr($c)):$c;
else {
for($j=0;$j<8&&(($c>>(7-$j))&1);$j++); // how many bytes represent this
Unicode character? $j
if(($j>=2 && $j<=6) && $len>$i+$j-1) { // valid range? is there $j-1
extra bytes?
$code = $c&pow(2, 7-$j)-1;
for($k=1;$k<$j;$k++) {
if(((ord($input{$i+$k})>>6)&3)==2)
$code = ($code<<6)|(ord($input{$i+$k})&63);
else {
$result['errorbytes'][]= $i;
continue 2;
}
}
$codepoints[]= $include_raw?array($code, substr($input, $i, $j)):$code;
$i+=$j-1;
}
else
$result['errorbytes'][]= $i;
}
$result['length']++;
}
$result['codepoints'] = $codepoints;
return($result);
}
?>
Some examples of the result:
<?php
$r = get_utf_8('è¿æ¥ã');
/*
Array(
[length] => 3
[errorbytes] => Array(
)
[codepoints] => Array(
[0] => 36814
[1] => 25509
[2] => 12290
)
)
*/
// Using the $include_raw parameter includes the raw UTF-8 bytes...
$r = get_utf_8('è¿æ¥ã', TRUE);
/*
Array(
[length] => 3
[errorbytes] => Array(
)
[codepoints] => Array(
[0] => Array(
[0] => 36814
[1] => è¿
)
[1] => Array(
[0] => 25509
[1] => æ¥
)
[2] => Array(
[0] => 12290
[1] => ã
)
)
)
*/
?>