Doc #80166 [Com]: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u
| From: | thomas at landauer dot at | Date: | Thu, 01 Oct 2020 12:31:38 +0000 |
| Subject: | Doc #80166 [Com]: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u | ||
| References: | 1 | Groups: | php.doc.bugs |
| Request: | Send a blank email to doc-bugs+get-17943@lists.php.net to get a copy of this message | ||
Edit report at https://bugs.php.net/bug.php?id=80166&edit=1
ID: 80166
Comment by: thomas at landauer dot at
Reported by: thomas at landauer dot at
Summary: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring
modifier u
Status: Closed
Type: Documentation Problem
Package: PCRE related
Operating System: Linux
PHP Version: 7.2.33
Assigned To: cmb
Block user comment: N
Private report: N
New Comment:
> The use of character offsets should be avoided wherever possible.
Why?
When using *byte* offsets to split a string, you might end up with an invalid string, due to some
"half"-characters.
Previous Comments:
------------------------------------------------------------------------
[2020-10-01 12:20:12] phpdocbot@php.net
Automatic comment on behalf of cmb
Revision: http://git.php.net/?p=doc/en.git;a=commit;h=7d4c08228e566afb3caedd29cb837a2fa67fcfbf
Log: Fix #80166: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u
------------------------------------------------------------------------
[2020-10-01 12:16:41] nikic@php.net
Right, the clarification here should be that this is always a "byte offset" including in
UTF8 mode.
------------------------------------------------------------------------
[2020-10-01 12:14:30] cmb@php.net
> PREG_OFFSET_CAPTURE always counts utf-8 special characters as
> *2* (bytes)
To clarify: UTF-8 characters are not necessarily counter as 2
bytes.
------------------------------------------------------------------------
[2020-10-01 12:10:26] nikic@php.net
The documentation should be clarified, but the current behavior is both intentional and preferable.
The use of character offsets should be avoided wherever possible.
------------------------------------------------------------------------
[2020-10-01 12:02:18] thomas at landauer dot at
Description:
------------
PREG_OFFSET_CAPTURE always counts utf-8 special characters as *2* (bytes) - no matter if the
modifier
u (PCRE_UTF8) is present or not.
This has been reported before at https://bugs.php.net/bug.php?id=37391 but got
closed as "Not a bug". However, taking a look at the documentation clearly shows that this
*is* a bug:
https://www.php.net/manual/en/function.preg-match-all.php#refsect1-function.preg-match-all-parameters
says about PREG_OFFSET_CAPTURE
> If this flag is passed, for every occurring match the appendant string offset will also be
> returned.
"string offset" obviously refers to the number of *characters* in front of the match.
And https://www.php.net/manual/en/reference.pcre.pattern.modifiers.php
says about u (PCRE_UTF8):
> Pattern and subject strings are treated as UTF-8.
My actual PHP version is 7.2.24
Test script:
---------------
preg_match_all('/a/u', 'öa', $matches, PREG_OFFSET_CAPTURE);
var_dump($matches);
Expected result:
----------------
[1]=>int(1)
Actual result:
--------------
[1]=>int(2)
(=expected result without modifier u)
------------------------------------------------------------------------
--
Edit this bug report at https://bugs.php.net/bug.php?id=80166&edit=1