[php-src] PR #24207: ext/standard: Skip mbrlen() for ASCII bytes in php_mblen()
| From: | ArtUkrainskiy | Date: | Thu, 08 Oct 2026 21:18:46 +0000 |
| Subject: | [php-src] PR #24207: ext/standard: Skip mbrlen() for ASCII bytes in php_mblen() | ||
| Groups: | php.git-pulls | ||
| Request: | Send a blank email to git-pulls+get-39273@lists.php.net to get a copy of this message | ||
Pull Request: https://github.com/php/php-src/pull/24207
Author: ArtUkrainskiy
fgetcsv(), str_getcsv(), escapeshellarg() and
escapeshellcmd() call php_mblen() while scanning their input; for ASCII
input this means one call per byte. With glibc this goes through mbrlen() into gconv,
accounting for about 80% of the instructions in a 10-column fgetcsv() loop (callgrind).
When Zend classifies the current locale as using ASCII characters as singletons
(CG(ascii_compatible_locale), the same flag php_basename() already relies
on), a non-NUL byte below 0x80 is one character and can be counted without calling
mbrlen().
php_mb_reset(), which callers run at the start of a string, records that flag;
escapeshellarg() and escapeshellcmd() now call it too, like
fgetcsv() and basename(). Other locales continue through
mbrlen() with conversion state preserved across the string, as ZTS builds already did;
NTS builds previously used mblen() and now share the same explicit-state path.
Some glibc converters (BIG5-HKSCS, CP1255, JIS X 0213, TSCII, ...) can flush a buffered character
without consuming the current input byte, causing mbrlen() to return 0 for
a non-NUL byte. In that case php_mblen() calls mbrlen() again with the
updated state and the same input byte, so the caller gets a consumed length and can make progress.
### Benchmark
Release build, ns per call:
| Benchmark | master | this |
|---|---:|---:|
| fgetcsv(), 10 columns | 4,217 | 515 |
| fgetcsv(), 50 columns | 20,775 | 2,241 |
| str_getcsv(), 10 columns | 4,131 | 458 |
| escapeshellarg(), 24-byte path | 310 | 48 |
| escapeshellarg(), 128 bytes, ja_JP.SJIS (old path) | 1,400 | 1,368 |
### Differential testing
I compared every caller against master NTS:
- str_getcsv() with three dialects
- fgetcsv()
- escapeshellarg()
- escapeshellcmd()
- basename()
- pathinfo()
The test used ~6,500 inputs mixing lead bytes, ASCII-special bytes in trail-byte positions, and lone
high bytes across 19 glibc locales, including C, UTF-8, Shift_JIS, GBK, GB18030, EUC-JP/KR/TW, Big5,
Big5-HKSCS, CP1255, JIS X 0213, TSCII and IBM1047, built with localedef.
16/19 locales produced identical output in both the new NTS and ZTS builds.
The three differences are explainable:
- **TCVN5712-1 and CP1258:** glibc's decoders can buffer an ASCII letter and combine it with a
following tone mark. On master NTS this can make e.g. a| appear as one multibyte
character, allowing | to pass through escapeshellcmd(). CP1258 is
classified by Zend as a single-byte locale, so the ASCII letter is now counted immediately and
| is escaped. For TCVN, NTS now preserves conversion state across the string, matching
the behavior the ZTS path already had.
- **IBM1047/EBCDIC:** this is single-byte and therefore falls under Zend's existing
ascii_compatible_locale classification, even though the encoding itself is not
ASCII-compatible (localedef warns about this). Bytes below 0x80 that glibc
previously reported as invalid are consequently counted as single bytes by the fast path.
ext/standard and ext/spl tests pass on both NTS and ZTS builds.
Benchmark and differential-test scripts:
https://github.com/ArtUkrainskiy/php-src-bench/tree/main/reports/fgetcsv-mblen