Bug #66121 [Com]: UTF-8 lookbehinds match bytes instead of characters

From: Date: Tue, 16 Dec 2014 11:01:26 +0000
Subject: Bug #66121 [Com]: UTF-8 lookbehinds match bytes instead of characters
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-189082@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=66121&edit=1

 ID:                 66121
 Comment by:         nhahtdh at gmail dot com
 Reported by:        danielklein at airpost dot net
 Summary:            UTF-8 lookbehinds match bytes instead of characters
 Status:             Open
 Type:               Bug
 Package:            PCRE related
 PHP Version:        5.5.6
 Block user comment: N
 Private report:     N

 New Comment:

In the "non-zero-width match and not reached limit" branch, if the startoffset is taken
from the match result, then the offset is advanced correctly, except for the case of using \C (the
behavior of PHP in such case is different from pcretest, but it should belong in a different bug
report).

In the "zero-width match or limit == 0" branch, except for the case in question, we break
out of the loop.

In the else branch, we also break out of the loop.

So I think everything should be fixed if we change this branch for all related functions.


Previous Comments:
------------------------------------------------------------------------
[2014-12-16 03:55:37] danielklein at airpost dot net

Same error seems to be on lines 838, 1296 & 1720. No idea if that would fix everything though.

------------------------------------------------------------------------
[2014-12-16 03:28:53] danielklein at airpost dot net

Good find but it looks like that code is just for pcre_replace(). It needs to be fixed for all
affected functions at the same time.

------------------------------------------------------------------------
[2014-12-15 11:19:49] nhahtdh at gmail dot com

The problem is most likely caused by advancing offset[1] (to which start_offset is assigned to and
used in the next call to pcre_exec) by 1 byte, regardless of UTF-8 mode or not:

https://github.com/php/php-src/blob/bf59acdea75cf13d179f10ce89d296a30f38676d/ext/pcre/php_pcre.c#L1296

The code is in the branch where we found out that we are standing still due to zero-length match,
and we can't find a non-zero-length match at the current position.

We need to check the actual mode and increment the offset by 1 UTF character if UTF-8 mode;
otherwise, 1 byte increment as per usual.

------------------------------------------------------------------------
[2013-11-22 04:14:54] danielklein at airpost dot net

The bug is also in preg_match_all(). See the following code:
<?php
preg_match_all('/(?<!ක)/u', 'ම', $matches, PREG_OFFSET_CAPTURE);
var_dump($matches);
?>

There should only be two matches, not four, (both empty strings) and the second one should have an
offset of 3 (bytes).

------------------------------------------------------------------------
[2013-11-20 23:31:30] danielklein at airpost dot net

I originally thought it was a PCRE bug so I submitted one with them (see http://bugs.exim.org/show_bug.cgi?id=1416 for
pcretest results). It has been confirmed to not be a PCRE bug so the bug must be in PHP.

------------------------------------------------------------------------


The remainder of the comments for this report are too long. To view
the rest of the comments, please view the bug report online at

    https://bugs.php.net/bug.php?id=66121


--
Edit this bug report at https://bugs.php.net/bug.php?id=66121&edit=1


Thread (11 messages)

« previous php.bugs (#189082) next »