Doc #78245 [Opn->Ver]: preg_split('/\R/', 'Техни') wrongly splits text into 2 parts

From: Date: Wed, 03 Jul 2019 13:59:45 +0000
Subject: Doc #78245 [Opn->Ver]: preg_split('/\R/', 'Техни') wrongly splits text into 2 parts
References: 1  Groups: php.doc.bugs 
Request: Send a blank email to doc-bugs+get-16792@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=78245&edit=1 ID: 78245 Updated by: cmb@php.net Reported by: bugs_php at zarevak dot net Summary: preg_split('/\R/', 'Техни') wrongly splits text into 2 parts -Status: Open +Status: Verified Type: Documentation Problem Package: PCRE related Operating System: any (tested Windows, Linux) PHP Version: 7.3.6 Block user comment: N Private report: N New Comment: From the PCRE2 docs[1]: | In 8-bit non-UTF-8 mode \R is equivalent to the following: | | (?>\r\n|\n|\x0b|\f|\r|\x85) [1] <https://www.pcre.org/current/doc/html/pcre2pattern.html#newlineseq> Previous Comments: ------------------------------------------------------------------------ [2019-07-03 13:54:18] nikic@php.net \R matches any Unicode line break. To only match \r, \n and \r\n the (*BSR_ANYCRLF) mode needs to be used. ------------------------------------------------------------------------ [2019-07-03 13:50:06] bugs_php at zarevak dot net Description: ------------ According to https://www.php.net/manual/en/regexp.reference.escape.php the \R should match line-breaks (\n, \r and \r\n), but here it matches something else and breaks the string into two. When using expression \r\n|\n|\r directly, it works correctly. Adding 'u' pattern modifier fixes the issue, but checks validity of the incoming UTF-8 string, which I do not want. The Russian text 'Техни' in this example does not contain \r or \n characters even when encoded using UTF-8 so this should not apply. The \R splits the string in the middle of the 'х' (U+0445 CYRILLIC SMALL LETTER HA) character, which is encoded in UTF-8 as bytes #D1 #85. This is either bug: 1] in implementation of \R, where it incorrectly matches different bytes then 13, 10 and their combination 2] or in documentation, where it should state what other characters it can match (in this example it seems, it matches U+0085 <control> : NEXT LINE [NEL]) Test script: --------------- <?php $string="Реконструкция\r\nРеконструкция - Служба Технической поддержки\r\n"; $array = preg_split('/\R/', $string); // BUG! $array2 = preg_split('/\r\n|\n|\r/', $string); //OK :) var_dump($array, $array2, ($array===$array2)); Expected result: ---------------- array(3) { [0]=> string(26) "Реконструкция" [1]=> string(83) "Реконструкция - Служба Технической поддержки" [2]=> string(0) "" } array(3) { [0]=> string(26) "Реконструкция" [1]=> string(83) "Реконструкция - Служба Технической поддержки" [2]=> string(0) "" } bool(true) Actual result: -------------- array(4) { [0]=> string(26) "Реконструкция" [1]=> string(47) "Реконструкция - Служба Те" [2]=> string(35) "нической поддержки" [3]=> string(0) "" } array(3) { [0]=> string(26) "Реконструкция" [1]=> string(83) "Реконструкция - Служба Технической поддержки" [2]=> string(0) "" } bool(false) ------------------------------------------------------------------------ -- Edit this bug report at https://bugs.php.net/bug.php?id=78245&edit=1

« previous php.doc.bugs (#16792) next »