Edit report at https://bugs.php.net/bug.php?id=66121&edit=1
ID: 66121
Updated by: rasmus@php.net
Reported by: danielklein at airpost dot net
Summary: UTF-8 lookbehinds match bytes instead of characters
-Status: Open
+Status: Feedback
Type: Bug
Package: PCRE related
PHP Version: 5.5.6
Block user comment: N
Private report: N
New Comment:
Please test this using the pcretest command line tool which is part of the PCRE package. If you can
reproduce the issue there too, file the bug with the PCRE project. If you can't, we'll
look into it further here.
Previous Comments:
------------------------------------------------------------------------
[2013-11-20 00:01:31] danielklein at airpost dot net
Description:
------------
The test script appears to check every byte position within the UTF-8 encoded character and
backtrack until a valid starting byte is found, then check it. If you remove the improperly inserted
characters the resulting bytes do encode the original character correctly.
Once it has matched the beginning of the string it should then move ahead by a character, not by a
byte.
Test script:
---------------
<?php
// Sinhala characters
print(preg_replace('/(?<!à¶)/u', '*', 'à¶') .
"\n"); // Works properly
print(preg_replace('/(?<!à¶)/u', '*', 'ම') .
"\n"); // Triggers the bug
// English characters
print(preg_replace('/(?<!k)/u', '*', 'k') . "\n"); //
Works properly
print(preg_replace('/(?<!k)/u', '*', 'm') . "\n"); //
Works properly
?>
Expected result:
----------------
*à¶
*ම*
*k
*m*
Actual result:
--------------
*à¶
*�*�*�*
*k
*m*
------------------------------------------------------------------------
--
Edit this bug report at https://bugs.php.net/bug.php?id=66121&edit=1