encode/decodeXmlEntities

From: Date: Tue, 04 Nov 2003 09:09:11 +0000
Subject: encode/decodeXmlEntities
Groups: php.pear.dev 
Request: Send a blank email to pear-dev+get-23222@lists.php.net to get a copy of this message
Begin Discussion Thread... Folks....If I'm not stupid (which is a big possibility)...then we have some work here...
    /**
    * Escape XML entities.
    *
    * @param   string  xml      Text string to escape.
    *
    * @return  string  xml
    * @access  public
    */
    function encodeXmlEntities($xml)
    {
       $xml = str_replace(array('ü', 'Ü', 'ö',
                                 'Ö', 'ä', 'Ä',
                                 'ß', '<', '>',
                                 '"', '\''
                                ),
                           array('&#252;', '&#220;', '&#246;',
                                 '&#214;', '&#228;', '&#196;',
                                  '&#223;', '&lt;', '&gt;',
                                  '&quot;', '&apos;'
                                ),
                           $xml
                          );
        $xml = preg_replace(array("/\&([a-z\d\#]+)\;/i",
                                  "/\&/",
                                  "/\#\|\|([a-z\d\#]+)\|\|\#/i",
"/([^a-zA-Z\d\s\<\>\&\;\.\:\=\"\-\/\%\?\!\'\(\)\[\]\{\}\$\#\+\,\@_])/e"
                                 ),
                            array("#||\\1||#",
                                  "&amp;",
                                  "&\\1;",
                                  "'&#'.ord('\\1').';'"
                                 ),
                            $xml
                           );
        return $xml;
    }
This is right out of the the XML/Tree/Node.php source.... And...it don't work....OK, OK, so it does...a little...but...it is broke...or..again..I'm stupid... The following code is used to test... <?php include("XML/Tree.php"); // instantiate object $tree = new XML_Tree(); // add the root element $root =& $tree->addRoot("test", chr(27).chr(29)."Crap"); $root->addChild("Note","Patient: Ty"); $root->addChild("Japanese","こんにちわ"); // print ingredients only echo $root->get(); ?> Ok, konichiwa is the only japanese word I know how to type...but...it suffices.... Here is the output...:::: <test>&#27;&#29;Crap <Note>Patient: Ty&#23;</Note> <Japanese>&#227;&#129;&#147;&#227;&#130;&#147;&#227;&#129;&#171;&#227;&#129;&#161;&#227;&#130;&#143;</Japanese> </test> Expected output.....:::: <test>&#27;&#29;Crap <Note>Patient: Ty&#23;</Note> <Japanese>こんにちわ</Japanese> </test> Oh...and the HTMLEntities created in the output....does NOT even decode correctly in a Browser...comes out as: Bad Output: Crap Patient: Ty ã�“ã‚“ã�«ã�¡ã‚� Expected Output: Crap Patient: Ty こんにちわ We must build an encode/decodeXmlEntities that recognizes valid UTF-8 characters (or any other characterset ISOxxx or EUCxxx) and if we continue to use perl regular expressions here...expand the regular expression to include only those characters that MUST be transformed. What we need is an table similar to HTML_ENTITIES that is XMLUTF8_ENTITIES that transform only INVALID XML characters.. Somehwere I read the following (paraphrased) "XML content is limited to valid HTML content"...well, one might assume that we have to use HTMLEntities or something similar... Well, that's wrong...we cannot. UTF-8 characterset is VALID hmtl I must NOT transalate こんにちわ to &#xxx; characters. The japanese character set (especially Unicode) is a valid HTML character set. We in the Western world (including Europe) cannot write everthing such that it breaks in Asia or any Double-Byte characterset part of the world (enough of my soap box...) The next steps are: 1) Define what XML entities must be translated to &#xxx; characters and what must remain as "readable" text 2) Define a new HTML Entities routine (or modify the existing routine) to recognize the character set (UTF-8 or Unicode) and use that function in encodeXMLEntities()... Things to note: HTMLEntities will give a DIFFERENT result when using NON-ASCII characters if the PHP Internal Encoding is set to different formats:
    1 - ASCII
    2 - EUC_JP
    3 - SJIS
    4 - UNICODE
    5 - other....
I don't like HTMLEntities...and have never used it. All browsers today will recognize all characters without the &#xxx; representation.... What we have to work on is XML. There are reasons XML cannot transmit "binary" and I say "binary" because ASCII IS binary..It's just Text...and it is Text...and XML is defined as "TEXT" only...and therefore characters that are NOT text must be translated to a text form... The difficulty here is identifying WHAT is TEXT???? I hate to break it to the Western world (and me included) that こんいちわ is TEXT...it's just not A-Z and a-z... I have to admit the XML specs are confusing in this area (maybe they're not and I'm stupid)...but...We can't be relying on BYTE-CODE to convert text that's not A-Z and a-z to some sort of &#xxx; representation that is broken.. In the example above: the <Japanese>&#227;&#129;&#147;&#227;&#130;&#147;&#227;&#129;&#171;&#227;&#129;&#161;&#227;&#130;&#143;</Japanese> text....the &#227;&#129 doesn't appar to be unicode or utf-8 for こ . The UTF-8 sequence would be \u3053\u3059 ... the Unicode character..well, I'd have to look that up... So...How do we fix? Developer feedback needed...let's finish this thread and make a good product!! mbstring may be a start...and use a CHARACTER safe function (not a byte function) to parse through the string and translate characters....こ is a character...even if it is 2 bytes... cheers! James

« previous php.pear.dev (#23222) next »