Req #72655 [Com]: How to speed up hashing between 40% and 50%
| From: | croverwnorene8 at googlemail dot com | Date: | Sat, 11 Mar 2023 06:26:09 +0000 |
| Subject: | Req #72655 [Com]: How to speed up hashing between 40% and 50% | ||
| References: | 1 | Groups: | php.bugs |
| Request: | Send a blank email to php-bugs+get-243893@lists.php.net to get a copy of this message | ||
Edit report at https://bugs.php.net/bug.php?id=72655&edit=1
ID: 72655
Comment by: croverwnorene8 at googlemail dot com
Reported by: simonhf at gmail dot com
Summary: How to speed up hashing between 40% and 50%
Status: Open
Type: Feature/Change Request
Package: Strings related
Operating System: Irrelevant
PHP Version: 7.0.9
Block user comment: N
Private report: N
New Comment:
Thanks for the information.. https://www.umrproviderportal.org/github.com
Previous Comments:
------------------------------------------------------------------------
[2016-07-22 22:50:08] simonhf at gmail dot com
Description:
------------
I saw that the hash function is a relatively old byte by byte hash function. Newer hash functions
like murmurhash3 work on multiple bytes at a time and are therefore faster for bigger strings
hashed.
I instrumented the existing hash function to capture each string hashed. When running php -v then
about 8000 strings were captured. They were all under 39 bytes long.
When racing the existing hash function against mmh3 then they are about the same, and the only
advantage is that mmh3 would go faster if the strings were longer... which they are not.
However, if I pad out the 8000 real world strings to the next 16 bytes and re-race then mmh3 is
about 40% to 50% faster than the existing hash algorithm.
Why does mmh3 do better this time? Because it works much faster on blocks of 16 bytes. It processes
any tail of up to 15 bytes, byte by byte still. By padding out the 8000 real world strings to the
next 16 bytes then mmh3 never had to run its byte by byte tail code and is therefore much faster.
This got me to thinking that if PHP internals were changed to pad all strings to the next 16 byte
boundary and use mmh3 instead then wouldn't this lead to a measurable run-time performance
increase in exchange for a slight run-time memory increase?
[1] https://github.com/php/php-src/blob/PHP-7.1/Zend/zend_string.h#L325
Test script:
---------------
Key,value store is a little faster, memory is a little bigger.
------------------------------------------------------------------------
--
Edit this bug report at https://bugs.php.net/bug.php?id=72655&edit=1