[php-src] PR #24142: ext/standard: Stop at the first non-name byte in html_entity_decode()
| From: | ArtUkrainskiy | Date: | Mon, 05 Oct 2026 18:02:24 +0000 |
| Subject: | [php-src] PR #24142: ext/standard: Stop at the first non-name byte in html_entity_decode() | ||
| Groups: | php.git-pulls | ||
| Request: | Send a blank email to git-pulls+get-39178@lists.php.net to get a copy of this message | ||
Pull Request: https://github.com/php/php-src/pull/24142
Author: ArtUkrainskiy
# ext/standard: Stop at the first non-name byte in html_entity_decode()
Follow-up to #18092.
That PR made
html_entity_decode() a lot faster on normal text, but strings with many
& became slower, as I noted in its description. The reason is the ;
search: for every & we call memchr() to look for a ; up
to 32 bytes ahead and then hash whatever is in between, even when it is clearly not an entity
(&&, & , &foo&bar).
@bukka suggested in the review of #18092 that a plain inline loop could beat memchr()
here, since entity names are short. I didn't try it back then. Now I did, and it works: scan
the name while it is letters and digits and stop at the first other byte. If that byte is not a
;, it is not an entity and there is nothing to hash. I also skip the
memchr() for & when the next byte already is one.
html_entity_decode(), 4 KB inputs, master vs this PR:
| input | |
|---|---|
| only & | 3.5× faster |
| only entities | 1.1× faster |
| 1–25% of the bytes in entities | same to 1.06× faster |
| jQuery source | 1.1–1.4× faster |
| jQuery minified | 1.4–1.7× faster |
| HTML pages, Markdown | same |
| ¶m=value URLs in text | up to 1.07× slower |
The string of & is back to where it was before #18092. The one case that gets a bit
slower is raw URL parameters in text (&logoColor=white): the name is scanned byte
by byte where a single memchr() call used to be enough. I think that is an acceptable
trade.
The output does not change: I compared it with master on 50k random inputs with different flags and
charsets, and ext/standard/tests/strings passes. htmlspecialchars_decode()
goes through the same code.