note 111815 deleted from function.utf8-decode by cmb

From: Date: Sun, 11 Dec 2016 11:31:52 +0000
Subject: note 111815 deleted from function.utf8-decode by cmb
References: 1  Groups: php.notes 
Request: Send a blank email to php-notes+get-208529@lists.php.net to get a copy of this message
Note Submitter: Nitrogen ---- Hi all. I have written this comprehensive UTF-8 decoder as I needed a way to split any UTF-8 string into nice and useable pieces of information, such as an array of code points that make up the input string regardless if it's UTF-8 or not. Here's what's good about this function: 1) It splits a UTF-8 encoded string into an array of code points. 2) It supports from 1 byte (7 bit) to 6 byte (31 bit) UTF-8 characters. 3) It checks the validity of the UTF-8 input string and provides an array of erroneous UTF-8 offsets. 4) Provides you with the character length of the UTF-8 string. <?php function get_utf_8($input, $include_raw = FALSE) { $result = array( 'length'=>0, 'errorbytes'=>array(), 'codepoints'=>array() ); $codepoints = array(); $len = strlen($input); for($i=0;$i<$len;$i++) { $c = ord($input{$i}); $msb = ($c>>7)&1; $bytes = 1; if(!$msb) $codepoints[]= $include_raw?array($c, chr($c)):$c; else { for($j=0;$j<8&&(($c>>(7-$j))&1);$j++); // how many bytes represent this Unicode character? $j if(($j>=2 && $j<=6) && $len>$i+$j-1) { // valid range? is there $j-1 extra bytes? $code = $c&pow(2, 7-$j)-1; for($k=1;$k<$j;$k++) { if(((ord($input{$i+$k})>>6)&3)==2) $code = ($code<<6)|(ord($input{$i+$k})&63); else { $result['errorbytes'][]= $i; continue 2; } } $codepoints[]= $include_raw?array($code, substr($input, $i, $j)):$code; $i+=$j-1; } else $result['errorbytes'][]= $i; } $result['length']++; } $result['codepoints'] = $codepoints; return($result); } ?> Some examples of the result: <?php $r = get_utf_8('迎接。'); /* Array( [length] => 3 [errorbytes] => Array( ) [codepoints] => Array( [0] => 36814 [1] => 25509 [2] => 12290 ) ) */ // Using the $include_raw parameter includes the raw UTF-8 bytes... $r = get_utf_8('迎接。', TRUE); /* Array( [length] => 3 [errorbytes] => Array( ) [codepoints] => Array( [0] => Array( [0] => 36814 [1] => 迎 ) [1] => Array( [0] => 25509 [1] => 接 ) [2] => Array( [0] => 12290 [1] => 。 ) ) ) */ ?>

« previous php.notes (#208529) next »