Bug #67487 [NEW]: PREG_SPLIT_OFFSET_CAPTURE and UTF-8

From: Date: Fri, 20 Jun 2014 10:00:30 +0000
Subject: Bug #67487 [NEW]: PREG_SPLIT_OFFSET_CAPTURE and UTF-8
Groups: php.bugs 
Request: Send a blank email to php-bugs+get-186272@lists.php.net to get a copy of this message
From:             alexandre at abrioux dot fr
Operating system: windows
PHP version:      5.5.13
Package:          PCRE related
Bug Type:         Bug
Bug description:PREG_SPLIT_OFFSET_CAPTURE and UTF-8

Description:
------------
I quote :

PREG_SPLIT_OFFSET_CAPTURE

     "[...] the return value [is] an array where every element is an
array consisting of the matched string at offset 0 and its string offset
into subject at offset 1."

The "string offset" is wrong when the subject string is encoded in
UTF-8.
Maybe the function is using strlen() internally, instead of using
mb_strlen() ?

Note that I use the "u" modifier in the regex.

Test script:
---------------
<?php
header("Content-Type: text/plain; charset=utf-8");
var_dump(preg_split('# #u', 'à é ù', 0, PREG_SPLIT_OFFSET_CAPTURE));
?>

Actual result:
--------------
array(3) {
  [0]=>
  array(2) {
    [0]=>
    string(2) "à"
    [1]=>
    int(0)
  }
  [1]=>
  array(2) {
    [0]=>
    string(2) "é"
    [1]=>
    int(3)
  }
  [2]=>
  array(2) {
    [0]=>
    string(2) "ù"
    [1]=>
    int(6)
  }
}

-- 
Edit bug report at https://bugs.php.net/bug.php?id=67487&edit=1
-- 



Thread (3 messages)

« previous php.bugs (#186272) next »