Req #77769 [Com]: html_entity_decode does not decode all HTML5 entities
| From: | stvar at yahoo dot com | Date: | Sat, 11 May 2019 21:12:36 +0000 |
| Subject: | Req #77769 [Com]: html_entity_decode does not decode all HTML5 entities | ||
| References: | 1 | Groups: | php.bugs |
| Request: | Send a blank email to php-bugs+get-220821@lists.php.net to get a copy of this message | ||
Edit report at https://bugs.php.net/bug.php?id=77769&edit=1
ID: 77769
Comment by: stvar at yahoo dot com
Reported by: cananian at wikimedia dot org
Summary: html_entity_decode does not decode all HTML5
entities
Status: Open
Type: Feature/Change Request
Package: Strings related
Operating System: n/a
PHP Version: 7.3.3
Block user comment: N
Private report: N
New Comment:
Dear maintainers,
It is quite possible and feasible to have a more flexible
'html_entity_decode' that handles properly the named char
references that, for historical reasons, are allowed to
not be terminated with semicolon [1].
To sustain my claim, I invite you to examine Html-Cref
[2] -- a project that I developed quite recently which
implements several named character reference *parsers*
based on tries instead of hash tables.
Upon bringing into Html-Cref's framework PHP's function
'resolve_named_entity_html' and an adapted hash table
'ent_ht_html5' (all these as per the patch file enclosed;
the size of 'ent_ht_html5' was preserved), the standalone
binary obtained 'html-cref' is about 19% bigger then the
one built with the 'etrie' parser library: 203K vs. 163K.
Upon measurements, the new function 'html_cref_php_parse'
in 'src/html_cref_php.c' runs 4% slower than the fastest
trie-based parser on a 64-bit Intel Core I5-3210M machine.
Sincerely,
Stefan Vargyas.
[1] 12.2 Parsing HTML documents:
12.2.5.73 Named character reference state
https://html.spec.whatwg.org/#named-character-reference-state
[2] Html-Cref: Fast HTML Character References Decoder
https://github.com/stvar/html-cref
Previous Comments:
------------------------------------------------------------------------
[2019-03-19 21:05:35] requinix@php.net
They are decoded for graceful handling. &foo is still a parse error.
------------------------------------------------------------------------
[2019-03-19 20:10:11] cananian at wikimedia dot org
Description:
------------
The latest HTML5 specs contain a number of "semicolon-less" entities which are decoded in
most circumstances. See the list at https://html.spec.whatwg.org/#named-character-references
(just the ones which don't end in a semicolon).
These are decoded *except* when found in an attribute and the letter after the entity is an equals
sign or an ASCII alphanumeric; see https://html.spec.whatwg.org/#named-character-reference-state
I propose two new option flags for html_entity_decode:
ENT_HTML5_NOATTRIBUTE -- decodes all the semicolon-less entities in addition to the other HTML5
entities
ENT_HTML5_ATTRIBUTE -- decodes semicolon-less entities except when they are followed by an equals
sign or ASCII alphanumeric
This would allow authors to easily decode these legacy semicolon-less entities in the same way a
browser would.
Test script:
---------------
In PHP:
$ psysh
Psy Shell v0.9.9 (PHP 7.3.2-3 â cli) by Justin Hileman
>>> html_entity_decode('&foo', ENT_HTML5)
=> "&foo"
In a browser web console:
>document.body.innerHTML="&foo"
"&foo"
> document.body.innerHTML
"&foo"
------------------------------------------------------------------------
--
Edit this bug report at https://bugs.php.net/bug.php?id=77769&edit=1