Doc #78245 [Opn->Ver]: preg_split('/\R/', 'Ð¢ÐµÑ Ð½Ð¸') wrongly splits text into 2 parts
| From: | cmb@php.net | Date: | Wed, 03 Jul 2019 13:59:45 +0000 |
| Subject: | Doc #78245 [Opn->Ver]: preg_split('/\R/', 'Ð¢ÐµÑ Ð½Ð¸') wrongly splits text into 2 parts | ||
| References: | 1 | Groups: | php.doc.bugs |
| Request: | Send a blank email to doc-bugs+get-16792@lists.php.net to get a copy of this message | ||
Edit report at https://bugs.php.net/bug.php?id=78245&edit=1
ID: 78245
Updated by: cmb@php.net
Reported by: bugs_php at zarevak dot net
Summary: preg_split('/\R/', 'ТеÑ
ни') wrongly splits
text
into 2 parts
-Status: Open
+Status: Verified
Type: Documentation Problem
Package: PCRE related
Operating System: any (tested Windows, Linux)
PHP Version: 7.3.6
Block user comment: N
Private report: N
New Comment:
From the PCRE2 docs[1]:
| In 8-bit non-UTF-8 mode \R is equivalent to the following:
|
| (?>\r\n|\n|\x0b|\f|\r|\x85)
[1] <https://www.pcre.org/current/doc/html/pcre2pattern.html#newlineseq>
Previous Comments:
------------------------------------------------------------------------
[2019-07-03 13:54:18] nikic@php.net
\R matches any Unicode line break. To only match \r, \n and \r\n the (*BSR_ANYCRLF) mode needs to be
used.
------------------------------------------------------------------------
[2019-07-03 13:50:06] bugs_php at zarevak dot net
Description:
------------
According to https://www.php.net/manual/en/regexp.reference.escape.php
the \R should match line-breaks (\n, \r and \r\n), but here it matches something else and breaks the
string into two. When using expression \r\n|\n|\r directly, it works correctly.
Adding 'u' pattern modifier fixes the issue, but checks validity of the incoming UTF-8
string, which I do not want. The Russian text 'ТеÑ
ни' in this example does
not contain \r or \n characters even when encoded using UTF-8 so this should not apply. The \R
splits the string in the middle of the 'Ñ
' (U+0445 CYRILLIC SMALL LETTER HA) character,
which is encoded in UTF-8 as bytes #D1 #85.
This is either bug:
1] in implementation of \R, where it incorrectly matches different bytes then 13, 10 and their
combination
2] or in documentation, where it should state what other characters it can match (in this example it
seems, it matches U+0085 <control> : NEXT LINE [NEL])
Test script:
---------------
<?php
$string="РеконÑÑÑÑкÑиÑ\r\nРеконÑÑÑÑкÑиÑ
- СлÑжба ТеÑ
ниÑеÑкой
поддеÑжки\r\n";
$array = preg_split('/\R/', $string); // BUG!
$array2 = preg_split('/\r\n|\n|\r/', $string); //OK :)
var_dump($array, $array2, ($array===$array2));
Expected result:
----------------
array(3) {
[0]=>
string(26) "РеконÑÑÑÑкÑиÑ"
[1]=>
string(83) "РеконÑÑÑÑкÑÐ¸Ñ - СлÑжба
ТеÑ
ниÑеÑкой поддеÑжки"
[2]=>
string(0) ""
}
array(3) {
[0]=>
string(26) "РеконÑÑÑÑкÑиÑ"
[1]=>
string(83) "РеконÑÑÑÑкÑÐ¸Ñ - СлÑжба
ТеÑ
ниÑеÑкой поддеÑжки"
[2]=>
string(0) ""
}
bool(true)
Actual result:
--------------
array(4) {
[0]=>
string(26) "РеконÑÑÑÑкÑиÑ"
[1]=>
string(47) "РеконÑÑÑÑкÑÐ¸Ñ - СлÑжба
Те"
[2]=>
string(35) "ниÑеÑкой поддеÑжки"
[3]=>
string(0) ""
}
array(3) {
[0]=>
string(26) "РеконÑÑÑÑкÑиÑ"
[1]=>
string(83) "РеконÑÑÑÑкÑÐ¸Ñ - СлÑжба
ТеÑ
ниÑеÑкой поддеÑжки"
[2]=>
string(0) ""
}
bool(false)
------------------------------------------------------------------------
--
Edit this bug report at https://bugs.php.net/bug.php?id=78245&edit=1