note 47483 deleted from function.preg-match-all by tomsommer
| From: | tomsommer@php.net | Date: | Wed, 17 Nov 2004 19:41:44 +0000 |
| Subject: | note 47483 deleted from function.preg-match-all by tomsommer | ||
| References: | 1 | Groups: | php.notes |
| Request: | Send a blank email to php-notes+get-80665@lists.php.net to get a copy of this message | ||
Note Submitter: rogthefrog at earthlink dot net
----
preg_match_all and substr_count have the unfortunate side effect of not being as greedy as they
could be in case your pattern overlaps itself. For example, if you're looking for :1:2: and
have several adjacent :1:2: in your string, every other match will be skipped because, after a
match, the search resumes at n+1 and therefore skips the initial : in the immediately adjacent
match. Consider:
:1:2:1:2:1:2:3:4:5:1:2:
There are 4 instances of :1:2: in this string. preg_match_all and substr_count will find the first,
third and fourth, but skip the second one at position 4, because once the one at position 0 is
matched, matching resumes at position 5 (0 + length of match), i.e. the "1" character, and
so it doesn't find the :1:2: at position 4.
One workaround is to scan the string as follows:
<?
$ln = ":1:2:1:2:1:2:3:4:5:1:2:";
$pattern = ":1:2:";
$theCt = 0;
$i = 0;
while($i < strlen($ln))
{
$found = strpos(substr($ln, $i), $pattern);
if ($found !== false)
{
++$theCt;
$i += $found + strlen($pattern - 1);
}
else
++$i;
}
print "manual scan found ".$theCt." instances\n";
print "substr_count found ".substr_count($ln, $pattern)." instances\n";
preg_match_all("/".addslashes($pattern)."/", $ln, $matches);
print "preg_match_all found ".count($matches[0])." instances\n";
?>