note 27776 deleted from function.utf8-decode by cmb
| From: | cmb@php.net | Date: | Sun, 11 Dec 2016 11:32:56 +0000 |
| Subject: | note 27776 deleted from function.utf8-decode by cmb | ||
| References: | 1 | Groups: | php.notes |
| Request: | Send a blank email to php-notes+get-208532@lists.php.net to get a copy of this message | ||
Note Submitter: mittag - -add- -marcmittag- -dot- -de
----
The above function does not work entirely correct. It comes to problems, if there is a leading
"=" in one of the two Strings it produces and glues out of the two bytes of the unicode
letter.
The following works:
<?php
//Convert Unicode to ASCII + Entities
$fp = fopen($DOCUMENT_ROOT."/your_unicode_text.txt", "r");
while ( !feof($fp) )
{ $string = fgets($fp, 1000);
$utf2html_string .= $string;
}
$string2 = $utf2html_string;
fclose ( $fp);
function utf2html ()
{
global $utf2html_string;
$utf2html_retstr = "";
for ($utf2html_p=0; $utf2html_p<strlen($utf2html_string); $utf2html_p++) {
$utf2html_c = substr ($utf2html_string, $utf2html_p, 1);
$utf2html_c1 = ord ($utf2html_c);
if ($utf2html_c1>>5 == 6) {// 110x xxxx, 110 prefix for 2 bytes unicode
$utf2html_p++;
$utf2html_t = substr ($utf2html_string, $utf2html_p, 1);
$utf2html_c2 = ord ($utf2html_t);
$utf2html_c1 &= 31; // remove the 3 bit two bytes prefix
$utf2html_c2 &= 63; // remove the 2 bit trailing byte prefix
$utf2html_c2 |= (($utf2html_c1 & 3) << 6); // last 2 bits of c1 become first 2 of c2
$utf2html_c1 >>= 2; // c1 shifts 2 to the right
$a = dechex($utf2html_c1);
$a = str_pad($a, 2, "0", STR_PAD_LEFT);
$b = dechex($utf2html_c2);
$b = str_pad($b, 2, "0", STR_PAD_LEFT);
$utf2html_n_neu = $a.$b;
$utf2html_n_neu_speicher = $utf2html_n_neu;
$utf2html_n_neu = "&#x".$utf2html_n_neu.";";
$utf2html_retstr .= $utf2html_n_neu;
}
else {
$utf2html_retstr .= $utf2html_c;
}
}
echo $utf2html_retstr;
//return $utf2html_retstr;
}
utf2html();
?>