Doc #80166 [Csd]: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u

From: Date: Sat, 03 Oct 2020 21:19:12 +0000
Subject: Doc #80166 [Csd]: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u
References: 1  Groups: php.doc.bugs 
Request: Send a blank email to doc-bugs+get-17954@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=80166&edit=1 ID: 80166 User updated by: thomas at landauer dot at Reported by: thomas at landauer dot at Summary: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u Status: Closed Type: Documentation Problem Package: PCRE related Operating System: Linux PHP Version: 7.2.33 Assigned To: cmb Block user comment: N Private report: N New Comment: Here's what I posted to the "php.internals" mailing list: https://news-web.php.net/php.internals/111983 Previous Comments: ------------------------------------------------------------------------ [2020-10-01 14:25:10] phpdocbot@php.net Automatic comment on behalf of mumumu Revision: http://git.php.net/?p=doc/ja.git;a=commit;h=326a1d72dcef26f2461a23cf3b4897fab41f3375 Log: Fix #80166: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u ------------------------------------------------------------------------ [2020-10-01 13:44:00] nikic@php.net > And what's the point in having a dedicated modifier for UTF-8, if it doesn't make a > difference in the end? > I think you should support this modifier here too, and leave the decision (byte vs. character > offsets) to the user. I mean: This is exactly the point of such a switch, isn't it? The modifier controls how the pattern and input string are interpreted. You want /[äöü]/ to match the characters ä, ö and ü, not their constituent bytes in the UTF-8 encoding. It has no relation at all to the meaning of offsets, which are always proper byte offsets. ------------------------------------------------------------------------ [2020-10-01 13:11:23] cmb@php.net First, all offsets in the PCRE extension are byte offsets. Changing that would be a massive BC break. Second, UTF-8 character offsets enforce sequential access to characters and substrings, while byte offsets allow random access, which is way faster. If you still feel strongly that this should be changed, please write a mail to the internals mailing list[1], since this bugtracker is not suitable for this kind of discussion. [1] <https://www.php.net/mailing-lists.php#internals> ------------------------------------------------------------------------ [2020-10-01 12:58:20] thomas at landauer dot at What I meant: My current code works with character numbers (mb_substr(), mb_strlen()). If I switched to byte numbers, I'd have to change this to substr() and strlen(); and if anything goes wrong there, I'm not just extracting some wrong characters, but rather completely *destroying* the entire string... So why is it preferable to work with byte offsets? And what's the point in having a dedicated modifier for UTF-8, if it doesn't make a difference in the end? I think you should support this modifier here too, and leave the decision (byte vs. character offsets) to the user. I mean: This is exactly the point of such a switch, isn't it? ------------------------------------------------------------------------ [2020-10-01 12:38:47] nikic@php.net @thomas: If your pattern matches at a character boundary, then of course the returned byte offset will also always be located at a character boundary. The only way you could end up with a byte offset that is not on a character boundary is if your pattern explicitly requests that by using a single code unit match (\C). If it does, the result would not even be representable with a character offset. ------------------------------------------------------------------------ The remainder of the comments for this report are too long. To view the rest of the comments, please view the bug report online at https://bugs.php.net/bug.php?id=80166 -- Edit this bug report at https://bugs.php.net/bug.php?id=80166&edit=1

« previous php.doc.bugs (#17954) next »