Bug #66121 [Opn]: UTF-8 lookbehinds match bytes instead of characters

From: Date: Tue, 16 Dec 2014 03:28:54 +0000
Subject: Bug #66121 [Opn]: UTF-8 lookbehinds match bytes instead of characters
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-189079@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=66121&edit=1

 ID:                 66121
 User updated by:    danielklein at airpost dot net
 Reported by:        danielklein at airpost dot net
 Summary:            UTF-8 lookbehinds match bytes instead of characters
 Status:             Open
 Type:               Bug
 Package:            PCRE related
 PHP Version:        5.5.6
 Block user comment: N
 Private report:     N

 New Comment:

Good find but it looks like that code is just for pcre_replace(). It needs to be fixed for all
affected functions at the same time.


Previous Comments:
------------------------------------------------------------------------
[2014-12-15 11:19:49] nhahtdh at gmail dot com

The problem is most likely caused by advancing offset[1] (to which start_offset is assigned to and
used in the next call to pcre_exec) by 1 byte, regardless of UTF-8 mode or not:

https://github.com/php/php-src/blob/bf59acdea75cf13d179f10ce89d296a30f38676d/ext/pcre/php_pcre.c#L1296

The code is in the branch where we found out that we are standing still due to zero-length match,
and we can't find a non-zero-length match at the current position.

We need to check the actual mode and increment the offset by 1 UTF character if UTF-8 mode;
otherwise, 1 byte increment as per usual.

------------------------------------------------------------------------
[2013-11-22 04:14:54] danielklein at airpost dot net

The bug is also in preg_match_all(). See the following code:
<?php
preg_match_all('/(?<!ක)/u', 'ම', $matches, PREG_OFFSET_CAPTURE);
var_dump($matches);
?>

There should only be two matches, not four, (both empty strings) and the second one should have an
offset of 3 (bytes).

------------------------------------------------------------------------
[2013-11-20 23:31:30] danielklein at airpost dot net

I originally thought it was a PCRE bug so I submitted one with them (see http://bugs.exim.org/show_bug.cgi?id=1416 for
pcretest results). It has been confirmed to not be a PCRE bug so the bug must be in PHP.

------------------------------------------------------------------------
[2013-11-20 03:07:01] rasmus@php.net

Please test this using the pcretest command line tool which is part of the PCRE package. If you can
reproduce the issue there too, file the bug with the PCRE project. If you can't, we'll
look into it further here.

------------------------------------------------------------------------
[2013-11-20 00:01:31] danielklein at airpost dot net

Description:
------------
The test script appears to check every byte position within the UTF-8 encoded character and
backtrack until a valid starting byte is found, then check it. If you remove the improperly inserted
characters the resulting bytes do encode the original character correctly.

Once it has matched the beginning of the string it should then move ahead by a character, not by a
byte.

Test script:
---------------
<?php
// Sinhala characters
print(preg_replace('/(?<!ක)/u', '*', 'ක') .
"\n"); // Works properly
print(preg_replace('/(?<!ක)/u', '*', 'ම') .
"\n"); // Triggers the bug
// English characters
print(preg_replace('/(?<!k)/u', '*', 'k') . "\n"); //
Works properly
print(preg_replace('/(?<!k)/u', '*', 'm') . "\n"); //
Works properly
?>


Expected result:
----------------
*ක
*ම*
*k
*m*

Actual result:
--------------
*ක
*�*�*�*
*k
*m*


------------------------------------------------------------------------



--
Edit this bug report at https://bugs.php.net/bug.php?id=66121&edit=1


Thread (11 messages)

« previous php.bugs (#189079) next »