Doc #80166 [Csd]: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u
| From: | thomas at landauer dot at | Date: | Sat, 03 Oct 2020 21:19:12 +0000 |
| Subject: | Doc #80166 [Csd]: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u | ||
| References: | 1 | Groups: | php.doc.bugs |
| Request: | Send a blank email to doc-bugs+get-17954@lists.php.net to get a copy of this message | ||
Edit report at https://bugs.php.net/bug.php?id=80166&edit=1
ID: 80166
User updated by: thomas at landauer dot at
Reported by: thomas at landauer dot at
Summary: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring
modifier u
Status: Closed
Type: Documentation Problem
Package: PCRE related
Operating System: Linux
PHP Version: 7.2.33
Assigned To: cmb
Block user comment: N
Private report: N
New Comment:
Here's what I posted to the "php.internals" mailing list: https://news-web.php.net/php.internals/111983
Previous Comments:
------------------------------------------------------------------------
[2020-10-01 14:25:10] phpdocbot@php.net
Automatic comment on behalf of mumumu
Revision: http://git.php.net/?p=doc/ja.git;a=commit;h=326a1d72dcef26f2461a23cf3b4897fab41f3375
Log: Fix #80166: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u
------------------------------------------------------------------------
[2020-10-01 13:44:00] nikic@php.net
> And what's the point in having a dedicated modifier for UTF-8, if it doesn't make a
> difference in the end?
> I think you should support this modifier here too, and leave the decision (byte vs. character
> offsets) to the user. I mean: This is exactly the point of such a switch, isn't it?
The modifier controls how the pattern and input string are interpreted. You want /[äöü]/ to
match the characters ä, ö and ü, not their constituent bytes in the UTF-8 encoding. It has
no relation at all to the meaning of offsets, which are always proper byte offsets.
------------------------------------------------------------------------
[2020-10-01 13:11:23] cmb@php.net
First, all offsets in the PCRE extension are byte offsets.
Changing that would be a massive BC break.
Second, UTF-8 character offsets enforce sequential access to
characters and substrings, while byte offsets allow random access,
which is way faster.
If you still feel strongly that this should be changed, please
write a mail to the internals mailing list[1], since this
bugtracker is not suitable for this kind of discussion.
[1] <https://www.php.net/mailing-lists.php#internals>
------------------------------------------------------------------------
[2020-10-01 12:58:20] thomas at landauer dot at
What I meant: My current code works with character numbers (
mb_substr(),
mb_strlen()). If I switched to byte numbers, I'd have to change this to
substr() and strlen(); and if anything goes wrong there, I'm not just
extracting some wrong characters, but rather completely *destroying* the entire string...
So why is it preferable to work with byte offsets?
And what's the point in having a dedicated modifier for UTF-8, if it doesn't make a
difference in the end?
I think you should support this modifier here too, and leave the decision (byte vs. character
offsets) to the user. I mean: This is exactly the point of such a switch, isn't it?
------------------------------------------------------------------------
[2020-10-01 12:38:47] nikic@php.net
@thomas: If your pattern matches at a character boundary, then of course the returned byte offset
will also always be located at a character boundary.
The only way you could end up with a byte offset that is not on a character boundary is if your
pattern explicitly requests that by using a single code unit match (\C). If it does, the result
would not even be representable with a character offset.
------------------------------------------------------------------------
The remainder of the comments for this report are too long. To view
the rest of the comments, please view the bug report online at
https://bugs.php.net/bug.php?id=80166
--
Edit this bug report at https://bugs.php.net/bug.php?id=80166&edit=1