Bug #67487 [Opn->Nab]: PREG_SPLIT_OFFSET_CAPTURE and UTF-8
| From: | johannes@php.net | Date: | Fri, 20 Jun 2014 13:10:34 +0000 |
| Subject: | Bug #67487 [Opn->Nab]: PREG_SPLIT_OFFSET_CAPTURE and UTF-8 | ||
| References: | 1 | Groups: | php.bugs |
| Request: | Send a blank email to php-bugs+get-186277@lists.php.net to get a copy of this message | ||
Edit report at https://bugs.php.net/bug.php?id=67487&edit=1
ID: 67487
Updated by: johannes@php.net
Reported by: alexandre at abrioux dot fr
Summary: PREG_SPLIT_OFFSET_CAPTURE and UTF-8
-Status: Open
+Status: Not a bug
Type: Bug
Package: PCRE related
Operating System: windows
PHP Version: 5.5.13
Block user comment: N
Private report: N
New Comment:
This is consistent with PHP strings. A PHP string is an array of bytes, not unicode characters.
Previous Comments:
------------------------------------------------------------------------
[2014-06-20 10:00:29] alexandre at abrioux dot fr
Description:
------------
I quote :
PREG_SPLIT_OFFSET_CAPTURE
"[...] the return value [is] an array where every element is an array consisting of the
matched string at offset 0 and its string offset into subject at offset 1."
The "string offset" is wrong when the subject string is encoded in UTF-8.
Maybe the function is using strlen() internally, instead of using mb_strlen() ?
Note that I use the "u" modifier in the regex.
Test script:
---------------
<?php
header("Content-Type: text/plain; charset=utf-8");
var_dump(preg_split('# #u', 'à é ù', 0, PREG_SPLIT_OFFSET_CAPTURE));
?>
Actual result:
--------------
array(3) {
[0]=>
array(2) {
[0]=>
string(2) "Ã "
[1]=>
int(0)
}
[1]=>
array(2) {
[0]=>
string(2) "é"
[1]=>
int(3)
}
[2]=>
array(2) {
[0]=>
string(2) "ù"
[1]=>
int(6)
}
}
------------------------------------------------------------------------
--
Edit this bug report at https://bugs.php.net/bug.php?id=67487&edit=1