Bug #66121 [Com]: UTF-8 lookbehinds match bytes instead of characters

From: Date: Mon, 15 Dec 2014 11:19:50 +0000
Subject: Bug #66121 [Com]: UTF-8 lookbehinds match bytes instead of characters
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-189070@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=66121&edit=1 ID: 66121 Comment by: nhahtdh at gmail dot com Reported by: danielklein at airpost dot net Summary: UTF-8 lookbehinds match bytes instead of characters Status: Open Type: Bug Package: PCRE related PHP Version: 5.5.6 Block user comment: N Private report: N New Comment: The problem is most likely caused by advancing offset[1] (to which start_offset is assigned to and used in the next call to pcre_exec) by 1 byte, regardless of UTF-8 mode or not: https://github.com/php/php-src/blob/bf59acdea75cf13d179f10ce89d296a30f38676d/ext/pcre/php_pcre.c#L1296 The code is in the branch where we found out that we are standing still due to zero-length match, and we can't find a non-zero-length match at the current position. We need to check the actual mode and increment the offset by 1 UTF character if UTF-8 mode; otherwise, 1 byte increment as per usual. Previous Comments: ------------------------------------------------------------------------ [2013-11-22 04:14:54] danielklein at airpost dot net The bug is also in preg_match_all(). See the following code: <?php preg_match_all('/(?<!ක)/u', 'ම', $matches, PREG_OFFSET_CAPTURE); var_dump($matches); ?> There should only be two matches, not four, (both empty strings) and the second one should have an offset of 3 (bytes). ------------------------------------------------------------------------ [2013-11-20 23:31:30] danielklein at airpost dot net I originally thought it was a PCRE bug so I submitted one with them (see http://bugs.exim.org/show_bug.cgi?id=1416 for pcretest results). It has been confirmed to not be a PCRE bug so the bug must be in PHP. ------------------------------------------------------------------------ [2013-11-20 03:07:01] rasmus@php.net Please test this using the pcretest command line tool which is part of the PCRE package. If you can reproduce the issue there too, file the bug with the PCRE project. If you can't, we'll look into it further here. ------------------------------------------------------------------------ [2013-11-20 00:01:31] danielklein at airpost dot net Description: ------------ The test script appears to check every byte position within the UTF-8 encoded character and backtrack until a valid starting byte is found, then check it. If you remove the improperly inserted characters the resulting bytes do encode the original character correctly. Once it has matched the beginning of the string it should then move ahead by a character, not by a byte. Test script: --------------- <?php // Sinhala characters print(preg_replace('/(?<!ක)/u', '*', 'ක') . "\n"); // Works properly print(preg_replace('/(?<!ක)/u', '*', 'ම') . "\n"); // Triggers the bug // English characters print(preg_replace('/(?<!k)/u', '*', 'k') . "\n"); // Works properly print(preg_replace('/(?<!k)/u', '*', 'm') . "\n"); // Works properly ?> Expected result: ---------------- *ක *ම* *k *m* Actual result: -------------- *ක *�*�*�* *k *m* ------------------------------------------------------------------------ -- Edit this bug report at https://bugs.php.net/bug.php?id=66121&edit=1

« previous php.bugs (#189070) next »