Edit report at https://bugs.php.net/bug.php?id=62360&edit=1
ID: 62360
Comment by: minktee at hotmail dot com
Reported by: danielklein at airpost dot net
Summary: Five valid PCRE escape sequences not documented
Status: Open
Type: Documentation Problem
Package: Documentation problem
PHP Version: Irrelevant
Block user comment: N
Private report: N
New Comment:
Document problem
Previous Comments:
------------------------------------------------------------------------
[2012-08-02 03:50:50] vovan-ve at yandex dot ru
Also recursive (?+n) and (?-n) are undocumented (PHP >= 5.2.4, PCRE >= 7.2). Back references
in form \g<n>, \g<-n>, \g'n', \g'-n' (PHP >= 5.2.7, PCRE >=
7.7) are not mentioned.
------------------------------------------------------------------------
[2012-08-02 03:40:29] vovan-ve at yandex dot ru
Also \N is undocumented. It was introduced in PCRE/8.10 (PHP/5.3.4?).
Also some types of conditions in conditional subpattern (?( are undocumented. All available
conditions are:
n
+n, -n (PCRE >= 7.2, PHP >= 5.2.4)
name (PCRE >= 6.7, PHP >= 5.2.0)
<name>, 'name' (PCRE >= 7.0, PHP >= 5.2.2)
R
Rn (PCRE >= 7.0, PHP >= 5.2.2)
R&name (PCRE >= 7.0, PHP >= 5.2.2)
?assertion
DEFINE (PCRE >= 7.0, PHP >= 5.2.2)
Also (?C)/(?Cn) is not documented and not implemented. But it works and does nothing (PCRE >=
4.0).
The subject of the bug is "escape sequences not documented". Should it be renemed to
"features not documented"?
------------------------------------------------------------------------
[2012-06-19 01:28:09] danielklein at airpost dot net
Description:
------------
---
From manual page: http://www.php.net/regexp.reference.escape
---
There are currently five valid PCRE escape sequences not documented on this page: \C, \R, \X, \g
& \k. Some of these are documented on other pages.
\g & \k - regexp.reference.back-references
\C - regexp.reference.dot
\R matches line break characters or combinations (see http://nikic.github.com/2011/12/10/PCRE-and-newlines.html)
\X matches Unicode graphemes (see http://www.regular-expressions.info/unicode.html)
Also, the description for \G is misleading. \G can also match in preg_match_all or preg_replace at
the point where the previous match stopped (see example script).
Test script:
---------------
preg_match_all('/\b([^\Wv]+)\s+/X', "In the beginning the universe was created",
$matches, PREG_PATTERN_ORDER, 1); // Matches four words not containing 'v', then followed
by a space
var_export($matches[1]);
print("\n");
preg_match_all('/\G\b([^\Wv]+)\s+/X', "In the beginning the universe was
created", $matches, PREG_PATTERN_ORDER, 1); // No matches
var_export($matches[1]);
print("\n");
preg_match_all('/\G\b([^\Wv]+)\s+/X', "In the beginning the universe was
created", $matches, PREG_PATTERN_ORDER, 3); // Able to match \b at offset 3, then two others.
No word can have 'v' in it. \G prevents skipping to next word
var_export($matches[1]);
Actual result:
--------------
array (
0 => 'the',
1 => 'beginning',
2 => 'the',
3 => 'was',
)
array (
)
array (
0 => 'the',
1 => 'beginning',
2 => 'the',
)
------------------------------------------------------------------------
--
Edit this bug report at https://bugs.php.net/bug.php?id=62360&edit=1