Doc #80166 [Csd]: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u
| From: | nikic@php.net | Date: | Thu, 01 Oct 2020 12:38:47 +0000 |
| Subject: | Doc #80166 [Csd]: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u | ||
| References: | 1 | Groups: | php.doc.bugs |
| Request: | Send a blank email to doc-bugs+get-17944@lists.php.net to get a copy of this message | ||
Edit report at https://bugs.php.net/bug.php?id=80166&edit=1
ID: 80166
Updated by: nikic@php.net
Reported by: thomas at landauer dot at
Summary: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring
modifier u
Status: Closed
Type: Documentation Problem
Package: PCRE related
Operating System: Linux
PHP Version: 7.2.33
Assigned To: cmb
Block user comment: N
Private report: N
New Comment:
@thomas: If your pattern matches at a character boundary, then of course the returned byte offset
will also always be located at a character boundary.
The only way you could end up with a byte offset that is not on a character boundary is if your
pattern explicitly requests that by using a single code unit match (\C). If it does, the result
would not even be representable with a character offset.
Previous Comments:
------------------------------------------------------------------------
[2020-10-01 12:31:38] thomas at landauer dot at
> The use of character offsets should be avoided wherever possible.
Why?
When using *byte* offsets to split a string, you might end up with an invalid string, due to some
"half"-characters.
------------------------------------------------------------------------
[2020-10-01 12:20:12] phpdocbot@php.net
Automatic comment on behalf of cmb
Revision: http://git.php.net/?p=doc/en.git;a=commit;h=7d4c08228e566afb3caedd29cb837a2fa67fcfbf
Log: Fix #80166: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u
------------------------------------------------------------------------
[2020-10-01 12:16:41] nikic@php.net
Right, the clarification here should be that this is always a "byte offset" including in
UTF8 mode.
------------------------------------------------------------------------
[2020-10-01 12:14:30] cmb@php.net
> PREG_OFFSET_CAPTURE always counts utf-8 special characters as
> *2* (bytes)
To clarify: UTF-8 characters are not necessarily counter as 2
bytes.
------------------------------------------------------------------------
[2020-10-01 12:10:26] nikic@php.net
The documentation should be clarified, but the current behavior is both intentional and preferable.
The use of character offsets should be avoided wherever possible.
------------------------------------------------------------------------
The remainder of the comments for this report are too long. To view
the rest of the comments, please view the bug report online at
https://bugs.php.net/bug.php?id=80166
--
Edit this bug report at https://bugs.php.net/bug.php?id=80166&edit=1