Re: encode/decodeXMLEntities
| From: | James Barwick | Date: | Wed, 05 Nov 2003 05:49:40 +0000 |
| Subject: | Re: encode/decodeXMLEntities | ||
| References: | 1 2 | Groups: | php.pear.dev php.pear.dev |
| Request: | Send a blank email to pear-dev+get-23243@lists.php.net to get a copy of this message | ||
OK...I was being stupid...however...because I'm stupid...I was unaware of ANOTHER problem...
Let's take the following codes:
€�‚ƒ„…†‡ˆ‰Š‹Œ�Ž��‘’“”•–—˜™š›œ�žŸ ¡¢£¤¥¦§¨©ª«¬®¯°±²³´µ¶·¸¹º»¼½¾¿ÀÁÂÃÄÅÆÇÈÉÊËÌÍÎÏÐÑÒÓÔÕÖרÙÚÛÜÝÞßàáâãäåæçèéêëìíîïðñòóôõö÷øùúûüý
Because I've got mbstring turned on...and the internal PHP encoding set for UTF-8...these characters are encoded to their UTF-8 equivalents...eg..there is no single ASCII code equivelent or direct map in the UTF-8 character set..
What does this mean...
we'll, the $encoder array below which loads characters chr(192), chr(193), etc. needs to have the UNICODE characters...not these ASCII characters...and the text sent in the $xml variable needs to be UNICODE,
not ascii....and, it will be unicode if all is done right and mbstring's encoding is all unicode.
However, this makes the encodeXMLEntities a little more difficult as the
CHARACTERS that are in the array to match to must be in the Backend Encoding format. Perhaps everything should be run through as unicode and the encoding format be translated to unicode for this replacement.
James Barwick wrote:
In a fit of inspiration this morning...the following code works!!!!/** * Escape XML entities. * * @param string xml Text string to escape. * * @return string xml * @access public */ function encodeXmlEntities($xml) { $encoder = array( chr(0) => '�',chr(1) => '',chr(2) => '',chr(3) => '',chr(4) => '',chr(5) => '',chr(6) => '',chr(7) => '',chr(8) => '',chr(9) => '	', chr(10) => '
',chr(11) => '',chr(12) => '',chr(13) => '
',chr(14) => '',chr(15) => '',chr(16) => '',chr(17) => '',chr(18) => '',chr(19) => '', chr(20) => '',chr(21) => '',chr(22) => '',chr(23) => '',chr(24) => '',chr(25) => '',chr(26) => '',chr(27) => '',chr(28) => '',chr(29) => '', chr(30) => '',chr(31) => '',chr(34) => '"',chr(38) => '&', chr(60) => '<',chr(62) => '>', chr(127) => '',chr(128) => '€',chr(129) => '', chr(130) => '‚',chr(131) => 'ƒ',chr(132) => '„',chr(133) => '…',chr(134) => '†',chr(135) => '‡',chr(136) => 'ˆ',chr(137) => '‰',chr(138) => 'Š',chr(139) => '‹', chr(140) => 'Œ',chr(141) => '',chr(142) => 'Ž',chr(143) => '',chr(144) => '',chr(145) => '‘',chr(146) => '’',chr(147) => '“',chr(148) => '”',chr(149) => '•', chr(150) => '–',chr(151) => '—',chr(152) => '˜',chr(153) => '™',chr(154) => 'š',chr(155) => '›',chr(156) => 'œ',chr(157) => '',chr(158) => 'ž',chr(159) => 'Ÿ', chr(160) => ' ',chr(161) => '¡',chr(162) => '¢',chr(163) => '£',chr(164) => '¤',chr(165) => '¥',chr(166) => '¦',chr(167) => '§',chr(168) => '¨',chr(169) => '©', chr(170) => 'ª',chr(171) => '«',chr(172) => '¬',chr(173) => '­',chr(174) => '®',chr(175) => '¯',chr(176) => '°',chr(177) => '±',chr(178) => '²',chr(179) => '³', chr(180) => '´',chr(181) => 'µ',chr(182) => '¶',chr(183) => '·',chr(184) => '¸',chr(185) => '¹',chr(186) => 'º',chr(187) => '»',chr(188) => '¼',chr(189) => '½', chr(190) => '¾',chr(191) => '¿',chr(192) => 'À',chr(193) => 'Á',chr(194) => 'Â',chr(195) => 'Ã',chr(196) => 'Ä',chr(197) => 'Å',chr(198) => 'Æ',chr(199) => 'Ç', chr(200) => 'È',chr(201) => 'É',chr(202) => 'Ê',chr(203) => 'Ë',chr(204) => 'Ì',chr(205) => 'Í',chr(206) => 'Î',chr(207) => 'Ï',chr(208) => 'Ð',chr(209) => 'Ñ', chr(210) => 'Ò',chr(211) => 'Ó',chr(212) => 'Ô',chr(213) => 'Õ',chr(214) => 'Ö',chr(215) => '×',chr(216) => 'Ø',chr(217) => 'Ù',chr(218) => 'Ú',chr(219) => 'Û', chr(220) => 'Ü',chr(221) => 'Ý',chr(222) => 'Þ',chr(223) => 'ß',chr(224) => 'à',chr(225) => 'á',chr(226) => 'â',chr(227) => 'ã',chr(228) => 'ä',chr(229) => 'å', chr(230) => 'æ',chr(231) => 'ç',chr(232) => 'è',chr(233) => 'é',chr(234) => 'ê',chr(235) => 'ë',chr(236) => 'ì',chr(237) => 'í',chr(238) => 'î',chr(239) => 'ï', chr(240) => 'ð',chr(241) => 'ñ',chr(242) => 'ò',chr(243) => 'ó',chr(244) => 'ô',chr(245) => 'õ',chr(246) => 'ö',chr(247) => '÷',chr(248) => 'ø',chr(249) => 'ù', chr(250) => 'ú',chr(251) => 'û',chr(252) => 'ü',chr(253) => 'ý',chr(254) => 'þ',chr(255) => 'ÿ');$newstr=""; for ($i=0; $i<strlen($xml); $i++) { // substring is character safe with mbstring overloading... $c=substr($xml,$i,1); if (array_key_exists($c,$encoder)) $newstr.=$encoder[$c]; else $newstr.=$c; } return $newstr; }This code should be implemented in XML/Tree/Node.php as soon as possible so that we can all work with UTF8!!!! Any critisism? Feedback? I hope the PEAR developers are reading this...perhaps an email directly to Christian will help.... James James Barwick wrote:Followup - This works for me: Modification of XML/Tree/Node.php function encodeXmlEntities($xml){ $chrs = chr(0).chr(1).chr(2).chr(3).chr(4).chr(5).chr(6).chr(7).chr(8).chr(9).chr(10).chr(11).chr(12).chr(13).chr(14).chr(15).chr(16).chr(17).chr(18).chr(19). chr(20).chr(21).chr(22).chr(23).chr(24).chr(25).chr(26).chr(27).chr(28).chr(29).chr(30).chr(31).chr(34).chr(38). chr(60).chr(62). chr(127).chr(128).chr(129).chr(130).chr(131).chr(132).chr(133).chr(134).chr(135).chr(136).chr(137).chr(138).chr(139). chr(140).chr(141).chr(142).chr(143).chr(144).chr(145).chr(146).chr(147).chr(148).chr(149). chr(150).chr(151).chr(152).chr(153).chr(154).chr(155).chr(156).chr(157).chr(158).chr(159). chr(160).chr(161).chr(162).chr(163).chr(164).chr(165).chr(166).chr(167).chr(168).chr(169). chr(170).chr(171).chr(172).chr(173).chr(174).chr(175).chr(176).chr(177).chr(178).chr(179). chr(180).chr(181).chr(182).chr(183).chr(184).chr(185).chr(186).chr(187).chr(188).chr(189). chr(190).chr(191).chr(192).chr(193).chr(194).chr(195).chr(196).chr(197).chr(198).chr(199). chr(200).chr(201).chr(202).chr(203).chr(204).chr(205).chr(206).chr(207).chr(208).chr(209). chr(210).chr(211).chr(212).chr(213).chr(214).chr(215).chr(216).chr(217).chr(218).chr(219). chr(220).chr(221).chr(222).chr(223).chr(224).chr(225).chr(226).chr(227).chr(228).chr(229). chr(230).chr(231).chr(232).chr(233).chr(234).chr(235).chr(236).chr(237).chr(238).chr(239). chr(240).chr(241).chr(242).chr(243).chr(244).chr(245).chr(246).chr(247).chr(248).chr(249). chr(250).chr(251).chr(252).chr(253).chr(254).chr(255); $repchrs=array('�','','','','','','','','','	', '
','','','
','','','','','','', '','','','','','','','','','','','','"','&', '<','>', '','€','','‚','ƒ','„','…','†','‡','ˆ','‰','Š','‹', 'Œ','','Ž','','','‘','’','“','”','•', '–','—','˜','™','š','›','œ','','ž','Ÿ', ' ','¡','¢','£','¤','¥','¦','§','¨','©', 'ª','«','¬','­','®','¯','°','±','²','³', '´','µ','¶','·','¸','¹','º','»','¼','½', '¾','¿','À','Á','Â','Ã','Ä','Å','Æ','Ç', 'È','É','Ê','Ë','Ì','Í','Î','Ï','Ð','Ñ', 'Ò','Ó','Ô','Õ','Ö','×','Ø','Ù','Ú','Û', 'Ü','Ý','Þ','ß','à','á','â','ã','ä','å', 'æ','ç','è','é','ê','ë','ì','í','î','ï', 'ð','ñ','ò','ó','ô','õ','ö','÷','ø','ù', 'ú','û','ü','ý','þ','ÿ');$newstr=""; for ($i=0; $i<strlen($xml); $i++) { // substring is character safe with mbstring overloading... $c=substr($xml,$i,1); $x=strpos($chrs,$c); // this is broken if the double-byte character is a sequence in the above array // any suggestions? can I check if only one byte long? how? if ($x===false) $newstr.=$c; else $newstr.=$repchrs[$x]; } return $newstr; }But I need help finishing... Any way to speed this up? Seems like a slow way to do it...but...may be the only way... And...need to fix a the bug noted above... I tried to be clever doing strpos...but...think I may need to change my mind and scan an array myself.. any help?