Edit report at https://bugs.php.net/bug.php?id=66121&edit=1
ID: 66121
User updated by: danielklein at airpost dot net
Reported by: danielklein at airpost dot net
Summary: UTF-8 lookbehinds match bytes instead of characters
Status: Open
Type: Bug
Package: PCRE related
PHP Version: 5.5.6
Block user comment: N
Private report: N
New Comment:
Good find but it looks like that code is just for pcre_replace(). It needs to be fixed for all
affected functions at the same time.
Previous Comments:
------------------------------------------------------------------------
[2014-12-15 11:19:49] nhahtdh at gmail dot com
The problem is most likely caused by advancing offset[1] (to which start_offset is assigned to and
used in the next call to pcre_exec) by 1 byte, regardless of UTF-8 mode or not:
https://github.com/php/php-src/blob/bf59acdea75cf13d179f10ce89d296a30f38676d/ext/pcre/php_pcre.c#L1296
The code is in the branch where we found out that we are standing still due to zero-length match,
and we can't find a non-zero-length match at the current position.
We need to check the actual mode and increment the offset by 1 UTF character if UTF-8 mode;
otherwise, 1 byte increment as per usual.
------------------------------------------------------------------------
[2013-11-22 04:14:54] danielklein at airpost dot net
The bug is also in preg_match_all(). See the following code:
<?php
preg_match_all('/(?<!à¶)/u', 'ම', $matches, PREG_OFFSET_CAPTURE);
var_dump($matches);
?>
There should only be two matches, not four, (both empty strings) and the second one should have an
offset of 3 (bytes).
------------------------------------------------------------------------
[2013-11-20 23:31:30] danielklein at airpost dot net
I originally thought it was a PCRE bug so I submitted one with them (see http://bugs.exim.org/show_bug.cgi?id=1416 for
pcretest results). It has been confirmed to not be a PCRE bug so the bug must be in PHP.
------------------------------------------------------------------------
[2013-11-20 03:07:01] rasmus@php.net
Please test this using the pcretest command line tool which is part of the PCRE package. If you can
reproduce the issue there too, file the bug with the PCRE project. If you can't, we'll
look into it further here.
------------------------------------------------------------------------
[2013-11-20 00:01:31] danielklein at airpost dot net
Description:
------------
The test script appears to check every byte position within the UTF-8 encoded character and
backtrack until a valid starting byte is found, then check it. If you remove the improperly inserted
characters the resulting bytes do encode the original character correctly.
Once it has matched the beginning of the string it should then move ahead by a character, not by a
byte.
Test script:
---------------
<?php
// Sinhala characters
print(preg_replace('/(?<!à¶)/u', '*', 'à¶') .
"\n"); // Works properly
print(preg_replace('/(?<!à¶)/u', '*', 'ම') .
"\n"); // Triggers the bug
// English characters
print(preg_replace('/(?<!k)/u', '*', 'k') . "\n"); //
Works properly
print(preg_replace('/(?<!k)/u', '*', 'm') . "\n"); //
Works properly
?>
Expected result:
----------------
*à¶
*ම*
*k
*m*
Actual result:
--------------
*à¶
*�*�*�*
*k
*m*
------------------------------------------------------------------------
--
Edit this bug report at https://bugs.php.net/bug.php?id=66121&edit=1