Req #79545 [Com]: mbstring functions are about 10x slower

From: Date: Tue, 23 Jun 2020 15:49:47 +0000
Subject: Req #79545 [Com]: mbstring functions are about 10x slower
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-227611@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=79545&edit=1

 ID:                 79545
 Comment by:         alexinbeijing at gmail dot com
 Reported by:        michael dot vorisek at emailc dot z
 Summary:            mbstring functions are about 10x slower
 Status:             Open
 Type:               Feature/Change Request
 Package:            Performance problem
 Operating System:   Linux
 PHP Version:        7.4.5
 Block user comment: N
 Private report:     N

 New Comment:

I have studied this problem a bit more and am looking at ways to speed up case conversion of
multi-byte strings.

Please note there is another reason why the plain strtolower is so much faster: It uses
SSE2 instructions to process blocks of 16 bytes at once. This accounts for a large part of the
difference in performance.


Previous Comments:
------------------------------------------------------------------------
[2020-06-09 20:40:50] alexinbeijing at gmail dot com

This is a tricky one.

It looks like there is nothing really crazy going on in mbstring which is causing this vast
disparity in performance.

mbstring bounces through a series of function calls to process *each byte* of the input string. None
of those function calls is doing a lot of work, but when you have 200,000,000 bytes to work on, the
overhead really adds up.

Getting significantly more performance out of it might require major redesign.

------------------------------------------------------------------------
[2020-04-30 11:00:56] michael dot vorisek at emailc dot z

Description:
------------
https://3v4l.org/bOdXP

Currently the performance difference between single-byte and multi-byte string function is about 5x
- 25x (yes, +400% - +2400%).

Please verify and comment.

If the performance can not be improved with mbstring, then checking if the input string does contain
only 0x00 - 0x7f chars will improve the overall/real PHP performace - check can be done very quickly
using modern/vectored CPU instructions - and if satisfied, use equivalent single-byte string
function instead of multi-byte one.

Test script:
---------------
$cnt = 100000;

$strs = [
    'empty' => '',
    'short' => 'zluty kun',
    'short_with_uc' => 'zluty Kun',
    'long' => str_repeat('this is about 10000 chars long string', 270),
    'long_with_uc' => str_repeat('this is about 10000 chars long String',
270),
    'short_utf8' => 'žlutý kůň',
    'short_utf8_with_uc' => 'Žlutý kŮň',
];

foreach ($strs as $k => $str) {
    $a1 = microtime(true);
    for($i=0; $i < $cnt; ++$i){
        $res = strtolower($str);
    }
    $t1 = microtime(true) - $a1;
    // echo 'it took ' . round($t1 * 1000, 3) . ' ms for ++$i'."\n";

    $a2 = microtime(true);
    for($i=0; $i < $cnt; $i++){
        $res = mb_strtolower($str);
    }
    $t2 = microtime(true) - $a2;
    // echo 'it took ' . round($t2 * 1000, 3) . ' ms for $i++'."\n";

    echo 'strtolower is '.round($t2/$t1, 2).'x faster than mb_strtolower for ' .
$k . "\n\n";
}


Expected result:
----------------
strtolower is 6.73x faster than mb_strtolower for empty

strtolower is 9.78x faster than mb_strtolower for short

strtolower is 7.13x faster than mb_strtolower for short_with_uc

strtolower is 24.88x faster than mb_strtolower for long

strtolower is 23.12x faster than mb_strtolower for long_with_uc

strtolower is 9.05x faster than mb_strtolower for short_utf8

strtolower is 9.93x faster than mb_strtolower for short_utf8_with_uc



Actual result:
--------------
Difference should be no larger than 1.5x at least for single-byte strings.


------------------------------------------------------------------------



--
Edit this bug report at https://bugs.php.net/bug.php?id=79545&edit=1


Thread (5 messages)

« previous php.bugs (#227611) next »