Bug #67487 [Nab]: PREG_SPLIT_OFFSET_CAPTURE and UTF-8

From: Date: Fri, 20 Jun 2014 13:28:33 +0000
Subject: Bug #67487 [Nab]: PREG_SPLIT_OFFSET_CAPTURE and UTF-8
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-186279@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=67487&edit=1 ID: 67487 User updated by: test at test dot com Reported by: test at test dot com Summary: PREG_SPLIT_OFFSET_CAPTURE and UTF-8 Status: Not a bug Type: Bug Package: PCRE related Operating System: windows PHP Version: 5.5.13 Block user comment: N Private report: N New Comment: . Previous Comments: ------------------------------------------------------------------------ [2014-06-20 13:10:33] johannes@php.net This is consistent with PHP strings. A PHP string is an array of bytes, not unicode characters. ------------------------------------------------------------------------ [2014-06-20 10:00:29] test at test dot com Description: ------------ I quote : PREG_SPLIT_OFFSET_CAPTURE "[...] the return value [is] an array where every element is an array consisting of the matched string at offset 0 and its string offset into subject at offset 1." The "string offset" is wrong when the subject string is encoded in UTF-8. Maybe the function is using strlen() internally, instead of using mb_strlen() ? Note that I use the "u" modifier in the regex. Test script: --------------- <?php header("Content-Type: text/plain; charset=utf-8"); var_dump(preg_split('# #u', 'à é ù', 0, PREG_SPLIT_OFFSET_CAPTURE)); ?> Actual result: -------------- array(3) { [0]=> array(2) { [0]=> string(2) "à" [1]=> int(0) } [1]=> array(2) { [0]=> string(2) "é" [1]=> int(3) } [2]=> array(2) { [0]=> string(2) "ù" [1]=> int(6) } } ------------------------------------------------------------------------ -- Edit this bug report at https://bugs.php.net/bug.php?id=67487&edit=1

« previous php.bugs (#186279) next »