Edit report at https://bugs.php.net/bug.php?id=66121&edit=1
ID: 66121
Comment by: nhahtdh at gmail dot com
Reported by: danielklein at airpost dot net
Summary: UTF-8 lookbehinds match bytes instead of characters
Status: Open
Type: Bug
Package: PCRE related
PHP Version: 5.5.6
Block user comment: N
Private report: N
New Comment:
In the "non-zero-width match and not reached limit" branch, if the startoffset is taken
from the match result, then the offset is advanced correctly, except for the case of using \C (the
behavior of PHP in such case is different from pcretest, but it should belong in a different bug
report).
In the "zero-width match or limit == 0" branch, except for the case in question, we break
out of the loop.
In the else branch, we also break out of the loop.
So I think everything should be fixed if we change this branch for all related functions.
Previous Comments:
------------------------------------------------------------------------
[2014-12-16 03:55:37] danielklein at airpost dot net
Same error seems to be on lines 838, 1296 & 1720. No idea if that would fix everything though.
------------------------------------------------------------------------
[2014-12-16 03:28:53] danielklein at airpost dot net
Good find but it looks like that code is just for pcre_replace(). It needs to be fixed for all
affected functions at the same time.
------------------------------------------------------------------------
[2014-12-15 11:19:49] nhahtdh at gmail dot com
The problem is most likely caused by advancing offset[1] (to which start_offset is assigned to and
used in the next call to pcre_exec) by 1 byte, regardless of UTF-8 mode or not:
https://github.com/php/php-src/blob/bf59acdea75cf13d179f10ce89d296a30f38676d/ext/pcre/php_pcre.c#L1296
The code is in the branch where we found out that we are standing still due to zero-length match,
and we can't find a non-zero-length match at the current position.
We need to check the actual mode and increment the offset by 1 UTF character if UTF-8 mode;
otherwise, 1 byte increment as per usual.
------------------------------------------------------------------------
[2013-11-22 04:14:54] danielklein at airpost dot net
The bug is also in preg_match_all(). See the following code:
<?php
preg_match_all('/(?<!à¶)/u', 'ම', $matches, PREG_OFFSET_CAPTURE);
var_dump($matches);
?>
There should only be two matches, not four, (both empty strings) and the second one should have an
offset of 3 (bytes).
------------------------------------------------------------------------
[2013-11-20 23:31:30] danielklein at airpost dot net
I originally thought it was a PCRE bug so I submitted one with them (see http://bugs.exim.org/show_bug.cgi?id=1416 for
pcretest results). It has been confirmed to not be a PCRE bug so the bug must be in PHP.
------------------------------------------------------------------------
The remainder of the comments for this report are too long. To view
the rest of the comments, please view the bug report online at
https://bugs.php.net/bug.php?id=66121
--
Edit this bug report at https://bugs.php.net/bug.php?id=66121&edit=1