Doc #80166 [Csd]: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u
| From: | cmb@php.net | Date: | Thu, 01 Oct 2020 13:11:23 +0000 |
| Subject: | Doc #80166 [Csd]: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u | ||
| References: | 1 | Groups: | php.doc.bugs |
| Request: | Send a blank email to doc-bugs+get-17947@lists.php.net to get a copy of this message | ||
Edit report at https://bugs.php.net/bug.php?id=80166&edit=1
ID: 80166
Updated by: cmb@php.net
Reported by: thomas at landauer dot at
Summary: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring
modifier u
Status: Closed
Type: Documentation Problem
Package: PCRE related
Operating System: Linux
PHP Version: 7.2.33
Assigned To: cmb
Block user comment: N
Private report: N
New Comment:
First, all offsets in the PCRE extension are byte offsets.
Changing that would be a massive BC break.
Second, UTF-8 character offsets enforce sequential access to
characters and substrings, while byte offsets allow random access,
which is way faster.
If you still feel strongly that this should be changed, please
write a mail to the internals mailing list[1], since this
bugtracker is not suitable for this kind of discussion.
[1] <https://www.php.net/mailing-lists.php#internals>
Previous Comments:
------------------------------------------------------------------------
[2020-10-01 12:58:20] thomas at landauer dot at
What I meant: My current code works with character numbers (
mb_substr(),
mb_strlen()). If I switched to byte numbers, I'd have to change this to
substr() and strlen(); and if anything goes wrong there, I'm not just
extracting some wrong characters, but rather completely *destroying* the entire string...
So why is it preferable to work with byte offsets?
And what's the point in having a dedicated modifier for UTF-8, if it doesn't make a
difference in the end?
I think you should support this modifier here too, and leave the decision (byte vs. character
offsets) to the user. I mean: This is exactly the point of such a switch, isn't it?
------------------------------------------------------------------------
[2020-10-01 12:38:47] nikic@php.net
@thomas: If your pattern matches at a character boundary, then of course the returned byte offset
will also always be located at a character boundary.
The only way you could end up with a byte offset that is not on a character boundary is if your
pattern explicitly requests that by using a single code unit match (\C). If it does, the result
would not even be representable with a character offset.
------------------------------------------------------------------------
[2020-10-01 12:31:38] thomas at landauer dot at
> The use of character offsets should be avoided wherever possible.
Why?
When using *byte* offsets to split a string, you might end up with an invalid string, due to some
"half"-characters.
------------------------------------------------------------------------
[2020-10-01 12:20:12] phpdocbot@php.net
Automatic comment on behalf of cmb
Revision: http://git.php.net/?p=doc/en.git;a=commit;h=7d4c08228e566afb3caedd29cb837a2fa67fcfbf
Log: Fix #80166: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u
------------------------------------------------------------------------
[2020-10-01 12:16:41] nikic@php.net
Right, the clarification here should be that this is always a "byte offset" including in
UTF8 mode.
------------------------------------------------------------------------
The remainder of the comments for this report are too long. To view
the rest of the comments, please view the bug report online at
https://bugs.php.net/bug.php?id=80166
--
Edit this bug report at https://bugs.php.net/bug.php?id=80166&edit=1