Req #77769 [Com]: html_entity_decode does not decode all HTML5 entities

From: Date: Sat, 11 May 2019 21:12:36 +0000
Subject: Req #77769 [Com]: html_entity_decode does not decode all HTML5 entities
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-220821@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=77769&edit=1 ID: 77769 Comment by: stvar at yahoo dot com Reported by: cananian at wikimedia dot org Summary: html_entity_decode does not decode all HTML5 entities Status: Open Type: Feature/Change Request Package: Strings related Operating System: n/a PHP Version: 7.3.3 Block user comment: N Private report: N New Comment: Dear maintainers, It is quite possible and feasible to have a more flexible 'html_entity_decode' that handles properly the named char references that, for historical reasons, are allowed to not be terminated with semicolon [1]. To sustain my claim, I invite you to examine Html-Cref [2] -- a project that I developed quite recently which implements several named character reference *parsers* based on tries instead of hash tables. Upon bringing into Html-Cref's framework PHP's function 'resolve_named_entity_html' and an adapted hash table 'ent_ht_html5' (all these as per the patch file enclosed; the size of 'ent_ht_html5' was preserved), the standalone binary obtained 'html-cref' is about 19% bigger then the one built with the 'etrie' parser library: 203K vs. 163K. Upon measurements, the new function 'html_cref_php_parse' in 'src/html_cref_php.c' runs 4% slower than the fastest trie-based parser on a 64-bit Intel Core I5-3210M machine. Sincerely, Stefan Vargyas. [1] 12.2 Parsing HTML documents: 12.2.5.73 Named character reference state https://html.spec.whatwg.org/#named-character-reference-state [2] Html-Cref: Fast HTML Character References Decoder https://github.com/stvar/html-cref Previous Comments: ------------------------------------------------------------------------ [2019-03-19 21:05:35] requinix@php.net They are decoded for graceful handling. &ampfoo is still a parse error. ------------------------------------------------------------------------ [2019-03-19 20:10:11] cananian at wikimedia dot org Description: ------------ The latest HTML5 specs contain a number of "semicolon-less" entities which are decoded in most circumstances. See the list at https://html.spec.whatwg.org/#named-character-references (just the ones which don't end in a semicolon). These are decoded *except* when found in an attribute and the letter after the entity is an equals sign or an ASCII alphanumeric; see https://html.spec.whatwg.org/#named-character-reference-state I propose two new option flags for html_entity_decode: ENT_HTML5_NOATTRIBUTE -- decodes all the semicolon-less entities in addition to the other HTML5 entities ENT_HTML5_ATTRIBUTE -- decodes semicolon-less entities except when they are followed by an equals sign or ASCII alphanumeric This would allow authors to easily decode these legacy semicolon-less entities in the same way a browser would. Test script: --------------- In PHP: $ psysh Psy Shell v0.9.9 (PHP 7.3.2-3 — cli) by Justin Hileman >>> html_entity_decode('&ampfoo', ENT_HTML5) => "&ampfoo" In a browser web console: >document.body.innerHTML="&ampfoo" "&ampfoo" > document.body.innerHTML "&foo" ------------------------------------------------------------------------ -- Edit this bug report at https://bugs.php.net/bug.php?id=77769&edit=1

« previous php.bugs (#220821) next »