Bug #66121 [Opn]: UTF-8 lookbehinds match bytes instead of characters

From: Date: Tue, 13 Jan 2015 08:35:31 +0000
Subject: Bug #66121 [Opn]: UTF-8 lookbehinds match bytes instead of characters
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-189926@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=66121&edit=1

 ID:                 66121
 User updated by:    danielklein at airpost dot net
 Reported by:        danielklein at airpost dot net
 Summary:            UTF-8 lookbehinds match bytes instead of characters
 Status:             Open
 Type:               Bug
 Package:            PCRE related
 PHP Version:        5.5.6
 Block user comment: N
 Private report:     N

 New Comment:

Although your test doesn't split the code point into bytes, it also does not give the correct
result. There should be a second asterisk after the Sinhala letter.


Previous Comments:
------------------------------------------------------------------------
[2015-01-08 22:24:51] cmbecker69 at gmx dot de

Rather interestingly, this seems to have worked before PHP
5.2.0: <http://3v4l.org/tL64Y>.

BTW: when talking about Unicode it might be best to avoid the term
"character". In this case it is about code points.

------------------------------------------------------------------------
[2014-12-16 11:01:25] nhahtdh at gmail dot com

In the "non-zero-width match and not reached limit" branch, if the startoffset is taken
from the match result, then the offset is advanced correctly, except for the case of using \C (the
behavior of PHP in such case is different from pcretest, but it should belong in a different bug
report).

In the "zero-width match or limit == 0" branch, except for the case in question, we break
out of the loop.

In the else branch, we also break out of the loop.

So I think everything should be fixed if we change this branch for all related functions.

------------------------------------------------------------------------
[2014-12-16 03:55:37] danielklein at airpost dot net

Same error seems to be on lines 838, 1296 & 1720. No idea if that would fix everything though.

------------------------------------------------------------------------
[2014-12-16 03:28:53] danielklein at airpost dot net

Good find but it looks like that code is just for pcre_replace(). It needs to be fixed for all
affected functions at the same time.

------------------------------------------------------------------------
[2014-12-15 11:19:49] nhahtdh at gmail dot com

The problem is most likely caused by advancing offset[1] (to which start_offset is assigned to and
used in the next call to pcre_exec) by 1 byte, regardless of UTF-8 mode or not:

https://github.com/php/php-src/blob/bf59acdea75cf13d179f10ce89d296a30f38676d/ext/pcre/php_pcre.c#L1296

The code is in the branch where we found out that we are standing still due to zero-length match,
and we can't find a non-zero-length match at the current position.

We need to check the actual mode and increment the offset by 1 UTF character if UTF-8 mode;
otherwise, 1 byte increment as per usual.

------------------------------------------------------------------------


The remainder of the comments for this report are too long. To view
the rest of the comments, please view the bug report online at

    https://bugs.php.net/bug.php?id=66121


--
Edit this bug report at https://bugs.php.net/bug.php?id=66121&edit=1


Thread (11 messages)

« previous php.bugs (#189926) next »