[php-src] PR #24207: ext/standard: Skip mbrlen() for ASCII bytes in php_mblen()

From: Date: Thu, 08 Oct 2026 21:18:46 +0000
Subject: [php-src] PR #24207: ext/standard: Skip mbrlen() for ASCII bytes in php_mblen()
Groups: php.git-pulls 
Request: Send a blank email to git-pulls+get-39273@lists.php.net to get a copy of this message
Pull Request: https://github.com/php/php-src/pull/24207 Author: ArtUkrainskiy fgetcsv(), str_getcsv(), escapeshellarg() and escapeshellcmd() call php_mblen() while scanning their input; for ASCII input this means one call per byte. With glibc this goes through mbrlen() into gconv, accounting for about 80% of the instructions in a 10-column fgetcsv() loop (callgrind). When Zend classifies the current locale as using ASCII characters as singletons (CG(ascii_compatible_locale), the same flag php_basename() already relies on), a non-NUL byte below 0x80 is one character and can be counted without calling mbrlen(). php_mb_reset(), which callers run at the start of a string, records that flag; escapeshellarg() and escapeshellcmd() now call it too, like fgetcsv() and basename(). Other locales continue through mbrlen() with conversion state preserved across the string, as ZTS builds already did; NTS builds previously used mblen() and now share the same explicit-state path. Some glibc converters (BIG5-HKSCS, CP1255, JIS X 0213, TSCII, ...) can flush a buffered character without consuming the current input byte, causing mbrlen() to return 0 for a non-NUL byte. In that case php_mblen() calls mbrlen() again with the updated state and the same input byte, so the caller gets a consumed length and can make progress. ### Benchmark Release build, ns per call: | Benchmark | master | this | |---|---:|---:| | fgetcsv(), 10 columns | 4,217 | 515 | | fgetcsv(), 50 columns | 20,775 | 2,241 | | str_getcsv(), 10 columns | 4,131 | 458 | | escapeshellarg(), 24-byte path | 310 | 48 | | escapeshellarg(), 128 bytes, ja_JP.SJIS (old path) | 1,400 | 1,368 | ### Differential testing I compared every caller against master NTS: - str_getcsv() with three dialects - fgetcsv() - escapeshellarg() - escapeshellcmd() - basename() - pathinfo() The test used ~6,500 inputs mixing lead bytes, ASCII-special bytes in trail-byte positions, and lone high bytes across 19 glibc locales, including C, UTF-8, Shift_JIS, GBK, GB18030, EUC-JP/KR/TW, Big5, Big5-HKSCS, CP1255, JIS X 0213, TSCII and IBM1047, built with localedef. 16/19 locales produced identical output in both the new NTS and ZTS builds. The three differences are explainable: - **TCVN5712-1 and CP1258:** glibc's decoders can buffer an ASCII letter and combine it with a following tone mark. On master NTS this can make e.g. a| appear as one multibyte character, allowing | to pass through escapeshellcmd(). CP1258 is classified by Zend as a single-byte locale, so the ASCII letter is now counted immediately and | is escaped. For TCVN, NTS now preserves conversion state across the string, matching the behavior the ZTS path already had. - **IBM1047/EBCDIC:** this is single-byte and therefore falls under Zend's existing ascii_compatible_locale classification, even though the encoding itself is not ASCII-compatible (localedef warns about this). Bytes below 0x80 that glibc previously reported as invalid are consequently counted as single bytes by the fast path. ext/standard and ext/spl tests pass on both NTS and ZTS builds. Benchmark and differential-test scripts: https://github.com/ArtUkrainskiy/php-src-bench/tree/main/reports/fgetcsv-mblen

« previous php.git-pulls (#39273) next »