#28654 [Asn->Fbk]: Possible bug in utf8_encode (bit operations)
| From: | moriyoshi@php.net | Date: | Mon, 12 Jul 2004 16:51:24 +0000 |
| Subject: | #28654 [Asn->Fbk]: Possible bug in utf8_encode (bit operations) | ||
| References: | 1 | Groups: | php.bugs |
| Request: | Send a blank email to php-bugs+get-62099@lists.php.net to get a copy of this message | ||
ID: 28654
Updated by: moriyoshi@php.net
Reported By: krausbn@php.net
-Status: Assigned
+Status: Feedback
Bug Type: *Languages/Translation
Operating System: WinXP
PHP Version: 4.3.4
Assigned To: moriyoshi
New Comment:
Most likely not a bug in PHP. The reporter better check
out things I pointed out.
Previous Comments:
------------------------------------------------------------------------
[2004-07-11 21:34:21] sniper@php.net
Moriyoshi: Was that last comment a statement of this being a bug in PHP
or what? Is this verified bug? Can you fix it if it is? (if it's not bug
-> bogus..)
------------------------------------------------------------------------
[2004-06-14 21:00:29] moriyoshi@php.net
Looks like you are trying to do the conversion between
the code page 1252 and UTF-8.
http://www.microsoft.com/globaldev/reference/sbcs/
1252.htm
Let alone mbstring, most of iconv() implementations
support CP1252 (a.k.a. IBM1252).
HTH
------------------------------------------------------------------------
[2004-06-10 00:11:14] krausbn@php.net
Hm, what ISO standard do I use (german, Win32) when I paste© Word
text into a textarea and post it to a PHP script?
Is it possible to solve my problem by converting my character encoding
to iso-8859-1 with the mb-functions?
------------------------------------------------------------------------
[2004-06-08 09:38:23] derick@php.net
utf8_encode only deals with iso-8859-1, which does not define
characters in the range from 128 to 160. Though it should probably just
replace those characters with a question mark, as that's how invalid
characters are usually converted.
------------------------------------------------------------------------
[2004-06-06 22:55:32] krausbn@php.net
Description:
------------
Hi!
I'm currently developing a nice script that generates OpenOffice SXW
files by filling the content.xml (which is UTF-8 encoded) with database
content. While trying to do this I found out that utf8_encode('')
(charcode 147) returns 'Â'. But when I checked the whole result in
OffenOffice '' is displayed as square (character unknown?!). So I made
some tests with UTF-8 conversion (even mb_* functions) and recognized
that characters between 128 and 160 returned by utf8_encode() dont
seem to match the standard. As mentioned above '' is returned as 'Â'
but should be 'â' (as you will get it using UltraEdit for
conversion).
Does anyone can give me some explanations here?
Im not familiar with this UTF-8 / bit-conversion stuff, but I dont
think PHP does what its supposed to do here. For a first workaround I
simply coded a custom_utf8_encode() that uses an own char map to
override this misbehaviour (see below). Can someone help my out with
this strange bug?!
Regards
Bjoern Kraus
function custom_utf8_encode($str)
{
$chrMap = array(128 => 'â', 129 => 'Â', 130 =>
'â', 131 =>
'Æ',
132 => 'â', 133 => 'â¦', 134 => 'â
', 135 =>
'â¡',
136 => 'Ë', 137 => 'â°', 138 => 'Å
', 139 =>
'â¹',
140 => 'Å', 141 => 'Â', 142 =>
'Ž', 143 =>
'Â',
144 => 'Â', 145 => 'â', 146 =>
'â', 147 =>
'â',
148 => 'â', 149 => 'â¢', 150 =>
'â', 151 =>
'â',
152 => 'Ë', 153 => 'â¢', 154 =>
'Å¡', 155 =>
'âº',
156 => 'Å', 157 => 'Â', 158 =>
'ž', 159 =>
'Ÿ');
$newStr = '';
for ($i = 0; $i < strlen($str); $i++) {
$chrVal = ord($str[$i]);
if ($chrVal > 127 && $chrVal < 160) {
$newStr .= $chrMap[$chrVal];
}
else {
$newStr .= utf8_encode($str[$i]);
}
}
return $newStr;
}
------------------------------------------------------------------------
--
Edit this bug report at http://bugs.php.net/?id=28654&edit=1