Doc #80166 [Csd]: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u

From: Date: Thu, 01 Oct 2020 12:58:20 +0000
Subject: Doc #80166 [Csd]: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u
References: 1  Groups: php.doc.bugs 
Request: Send a blank email to doc-bugs+get-17945@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=80166&edit=1 ID: 80166 User updated by: thomas at landauer dot at Reported by: thomas at landauer dot at Summary: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u Status: Closed Type: Documentation Problem Package: PCRE related Operating System: Linux PHP Version: 7.2.33 Assigned To: cmb Block user comment: N Private report: N New Comment: What I meant: My current code works with character numbers (mb_substr(), mb_strlen()). If I switched to byte numbers, I'd have to change this to substr() and strlen(); and if anything goes wrong there, I'm not just extracting some wrong characters, but rather completely *destroying* the entire string... So why is it preferable to work with byte offsets? And what's the point in having a dedicated modifier for UTF-8, if it doesn't make a difference in the end? I think you should support this modifier here too, and leave the decision (byte vs. character offsets) to the user. I mean: This is exactly the point of such a switch, isn't it? Previous Comments: ------------------------------------------------------------------------ [2020-10-01 12:38:47] nikic@php.net @thomas: If your pattern matches at a character boundary, then of course the returned byte offset will also always be located at a character boundary. The only way you could end up with a byte offset that is not on a character boundary is if your pattern explicitly requests that by using a single code unit match (\C). If it does, the result would not even be representable with a character offset. ------------------------------------------------------------------------ [2020-10-01 12:31:38] thomas at landauer dot at > The use of character offsets should be avoided wherever possible. Why? When using *byte* offsets to split a string, you might end up with an invalid string, due to some "half"-characters. ------------------------------------------------------------------------ [2020-10-01 12:20:12] phpdocbot@php.net Automatic comment on behalf of cmb Revision: http://git.php.net/?p=doc/en.git;a=commit;h=7d4c08228e566afb3caedd29cb837a2fa67fcfbf Log: Fix #80166: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u ------------------------------------------------------------------------ [2020-10-01 12:16:41] nikic@php.net Right, the clarification here should be that this is always a "byte offset" including in UTF8 mode. ------------------------------------------------------------------------ [2020-10-01 12:14:30] cmb@php.net > PREG_OFFSET_CAPTURE always counts utf-8 special characters as > *2* (bytes) To clarify: UTF-8 characters are not necessarily counter as 2 bytes. ------------------------------------------------------------------------ The remainder of the comments for this report are too long. To view the rest of the comments, please view the bug report online at https://bugs.php.net/bug.php?id=80166 -- Edit this bug report at https://bugs.php.net/bug.php?id=80166&edit=1

« previous php.doc.bugs (#17945) next »