Doc #80166 [Csd]: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u

From: Date: Thu, 01 Oct 2020 13:11:23 +0000
Subject: Doc #80166 [Csd]: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u
References: 1  Groups: php.doc.bugs 
Request: Send a blank email to doc-bugs+get-17947@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=80166&edit=1 ID: 80166 Updated by: cmb@php.net Reported by: thomas at landauer dot at Summary: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u Status: Closed Type: Documentation Problem Package: PCRE related Operating System: Linux PHP Version: 7.2.33 Assigned To: cmb Block user comment: N Private report: N New Comment: First, all offsets in the PCRE extension are byte offsets. Changing that would be a massive BC break. Second, UTF-8 character offsets enforce sequential access to characters and substrings, while byte offsets allow random access, which is way faster. If you still feel strongly that this should be changed, please write a mail to the internals mailing list[1], since this bugtracker is not suitable for this kind of discussion. [1] <https://www.php.net/mailing-lists.php#internals> Previous Comments: ------------------------------------------------------------------------ [2020-10-01 12:58:20] thomas at landauer dot at What I meant: My current code works with character numbers (mb_substr(), mb_strlen()). If I switched to byte numbers, I'd have to change this to substr() and strlen(); and if anything goes wrong there, I'm not just extracting some wrong characters, but rather completely *destroying* the entire string... So why is it preferable to work with byte offsets? And what's the point in having a dedicated modifier for UTF-8, if it doesn't make a difference in the end? I think you should support this modifier here too, and leave the decision (byte vs. character offsets) to the user. I mean: This is exactly the point of such a switch, isn't it? ------------------------------------------------------------------------ [2020-10-01 12:38:47] nikic@php.net @thomas: If your pattern matches at a character boundary, then of course the returned byte offset will also always be located at a character boundary. The only way you could end up with a byte offset that is not on a character boundary is if your pattern explicitly requests that by using a single code unit match (\C). If it does, the result would not even be representable with a character offset. ------------------------------------------------------------------------ [2020-10-01 12:31:38] thomas at landauer dot at > The use of character offsets should be avoided wherever possible. Why? When using *byte* offsets to split a string, you might end up with an invalid string, due to some "half"-characters. ------------------------------------------------------------------------ [2020-10-01 12:20:12] phpdocbot@php.net Automatic comment on behalf of cmb Revision: http://git.php.net/?p=doc/en.git;a=commit;h=7d4c08228e566afb3caedd29cb837a2fa67fcfbf Log: Fix #80166: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u ------------------------------------------------------------------------ [2020-10-01 12:16:41] nikic@php.net Right, the clarification here should be that this is always a "byte offset" including in UTF8 mode. ------------------------------------------------------------------------ The remainder of the comments for this report are too long. To view the rest of the comments, please view the bug report online at https://bugs.php.net/bug.php?id=80166 -- Edit this bug report at https://bugs.php.net/bug.php?id=80166&edit=1

« previous php.doc.bugs (#17947) next »