Doc #64260 [Wfx]: Documentation misleading
| From: | adam at acdinternet dot com | Date: | Wed, 27 Feb 2013 10:33:34 +0000 |
| Subject: | Doc #64260 [Wfx]: Documentation misleading | ||
| References: | 1 | Groups: | php.doc.bugs |
| Request: | Send a blank email to doc-bugs+get-9597@lists.php.net to get a copy of this message | ||
Edit report at https://bugs.php.net/bug.php?id=64260&edit=1
ID: 64260
User updated by: adam at acdinternet dot com
Reported by: adam at acdinternet dot com
Summary: Documentation misleading
Status: Wont fix
Type: Documentation Problem
Package: Documentation problem
Operating System: Ubuntu 10.4
PHP Version: 5.3.21
Block user comment: N
Private report: N
New Comment:
Yes OK.
Further research made me doubt my original assertions. It's never clear from the web regex
testers whether they have the PCRE modules installed/enabled, so I need to check it all on my own
system.
I still think there's some basic ambiguity, which would be helped by a list somewhere of
exactly which Unicode character ranges fall within [Common] - and the others. Otherwise it's
all down to trial and error?
But thanks very much for looking into this.
Previous Comments:
------------------------------------------------------------------------
[2013-02-27 04:37:10] frozenfire@php.net
This seems more like a data source bug than a documentation bug. The text "Those
that are not part of an identified script are lumped together as Common." is
lifted directly from the pcre manpage (http://pcre.org/pcre.txt). I think that
the documentation is intended to be correct in theory, but upstream may have a
bug that causes it not to be correct.
I would recommend filing a bug upstream, if the issue is one that is likely to
cause others problems in the future.
------------------------------------------------------------------------
[2013-02-21 00:48:32] adam at acdinternet dot com
Description:
------------
---
From manual page: http://www.php.net/regexp.reference.unicode
---
This page states "Those that are not part of an identified script are lumped together as
Common."
That means mutual exclusivity between the list below in the docs and 'Common'. In practise
this appears not to be true.
The character 'm' is in both 'Common' and 'Ogham' according to http://www.pagecolumn.com/tool/pregtest.htm
Evidence..
Try preg pattern:
/[\p{Ogham}]/iu
with Testing subject
moo
Then try
/[\p{Common}]/iu
..you get results (matches) both times. The documentation is therefore logically incorrect?
The character 'o' seems also to be in \p{Tifinagh}
This has caused me real problems, so bears correcting if I'm right.
------------------------------------------------------------------------
--
Edit this bug report at https://bugs.php.net/bug.php?id=64260&edit=1