Doc #80166 [Csd]: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u

From: Date: Thu, 01 Oct 2020 12:38:47 +0000
Subject: Doc #80166 [Csd]: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u
References: 1  Groups: php.doc.bugs 
Request: Send a blank email to doc-bugs+get-17944@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=80166&edit=1 ID: 80166 Updated by: nikic@php.net Reported by: thomas at landauer dot at Summary: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u Status: Closed Type: Documentation Problem Package: PCRE related Operating System: Linux PHP Version: 7.2.33 Assigned To: cmb Block user comment: N Private report: N New Comment: @thomas: If your pattern matches at a character boundary, then of course the returned byte offset will also always be located at a character boundary. The only way you could end up with a byte offset that is not on a character boundary is if your pattern explicitly requests that by using a single code unit match (\C). If it does, the result would not even be representable with a character offset. Previous Comments: ------------------------------------------------------------------------ [2020-10-01 12:31:38] thomas at landauer dot at > The use of character offsets should be avoided wherever possible. Why? When using *byte* offsets to split a string, you might end up with an invalid string, due to some "half"-characters. ------------------------------------------------------------------------ [2020-10-01 12:20:12] phpdocbot@php.net Automatic comment on behalf of cmb Revision: http://git.php.net/?p=doc/en.git;a=commit;h=7d4c08228e566afb3caedd29cb837a2fa67fcfbf Log: Fix #80166: preg_match_all()'s PREG_OFFSET_CAPTURE ignoring modifier u ------------------------------------------------------------------------ [2020-10-01 12:16:41] nikic@php.net Right, the clarification here should be that this is always a "byte offset" including in UTF8 mode. ------------------------------------------------------------------------ [2020-10-01 12:14:30] cmb@php.net > PREG_OFFSET_CAPTURE always counts utf-8 special characters as > *2* (bytes) To clarify: UTF-8 characters are not necessarily counter as 2 bytes. ------------------------------------------------------------------------ [2020-10-01 12:10:26] nikic@php.net The documentation should be clarified, but the current behavior is both intentional and preferable. The use of character offsets should be avoided wherever possible. ------------------------------------------------------------------------ The remainder of the comments for this report are too long. To view the rest of the comments, please view the bug report online at https://bugs.php.net/bug.php?id=80166 -- Edit this bug report at https://bugs.php.net/bug.php?id=80166&edit=1

« previous php.doc.bugs (#17944) next »