Bug #67487 [NEW]: PREG_SPLIT_OFFSET_CAPTURE and UTF-8
From: alexandre at abrioux dot fr
Operating system: windows
PHP version: 5.5.13
Package: PCRE related
Bug Type: Bug
Bug description:PREG_SPLIT_OFFSET_CAPTURE and UTF-8
Description:
------------
I quote :
PREG_SPLIT_OFFSET_CAPTURE
"[...] the return value [is] an array where every element is an
array consisting of the matched string at offset 0 and its string offset
into subject at offset 1."
The "string offset" is wrong when the subject string is encoded in
UTF-8.
Maybe the function is using strlen() internally, instead of using
mb_strlen() ?
Note that I use the "u" modifier in the regex.
Test script:
---------------
<?php
header("Content-Type: text/plain; charset=utf-8");
var_dump(preg_split('# #u', 'à é ù', 0, PREG_SPLIT_OFFSET_CAPTURE));
?>
Actual result:
--------------
array(3) {
[0]=>
array(2) {
[0]=>
string(2) "Ã "
[1]=>
int(0)
}
[1]=>
array(2) {
[0]=>
string(2) "é"
[1]=>
int(3)
}
[2]=>
array(2) {
[0]=>
string(2) "ù"
[1]=>
int(6)
}
}
--
Edit bug report at https://bugs.php.net/bug.php?id=67487&edit=1
--
Thread (3 messages)
- alexandre at abrioux dot fr