note 78032 deleted from function.ord by danbrown
| From: | danbrown@php.net | Date: | Tue, 11 Oct 2011 16:06:42 +0000 |
| Subject: | note 78032 deleted from function.ord by danbrown | ||
| References: | 1 | Groups: | php.notes |
| Request: | Send a blank email to php-notes+get-183796@lists.php.net to get a copy of this message | ||
Note Submitter: kerry at shetline dot com
----
Here's my take on an earlier-posted UTF-8 version of ord, suitable for iterating through a
string by Unicode value. The function can optionally take an index into a string, and optionally
return the number of bytes consumed by a character so that you know how much to increment the index
to get to the next character.
<?php
function ordUTF8($c, $index = 0, &$bytes = null)
{
$len = strlen($c);
$bytes = 0;
if ($index >= $len)
return false;
$h = ord($c{$index});
if ($h <= 0x7F) {
$bytes = 1;
return $h;
}
else if ($h < 0xC2)
return false;
else if ($h <= 0xDF && $index < $len - 1) {
$bytes = 2;
return ($h & 0x1F) << 6 | (ord($c{$index + 1}) & 0x3F);
}
else if ($h <= 0xEF && $index < $len - 2) {
$bytes = 3;
return ($h & 0x0F) << 12 | (ord($c{$index + 1}) & 0x3F) << 6
| (ord($c{$index + 2}) & 0x3F);
}
else if ($h <= 0xF4 && $index < $len - 3) {
$bytes = 4;
return ($h & 0x0F) << 18 | (ord($c{$index + 1}) & 0x3F) << 12
| (ord($c{$index + 2}) & 0x3F) << 6
| (ord($c{$index + 3}) & 0x3F);
}
else
return false;
}
?>