Doc #54154 [Opn->Csd]: Missing info about using \p{xx} and \P{xx} escape sequences with script names.

From: Date: Mon, 12 Nov 2012 10:01:20 +0000
Subject: Doc #54154 [Opn->Csd]: Missing info about using \p{xx} and \P{xx} escape sequences with script names.
References: 1  Groups: php.doc.bugs 
Request: Send a blank email to doc-bugs+get-9097@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=54154&edit=1 ID: 54154 Updated by: aharvey@php.net Reported by: hayk@php.net Summary: Missing info about using \p{xx} and \P{xx} escape sequences with script names. -Status: Open +Status: Closed Type: Documentation Problem Package: Documentation problem PHP Version: Irrelevant -Assigned To: +Assigned To: aharvey Block user comment: N Private report: N New Comment: This bug has been fixed in the documentation's XML sources. Since the online and downloadable versions of the documentation need some time to get updated, we would like to ask you to be a bit patient. Thank you for the report, and for helping us make our documentation better. Previous Comments: ------------------------------------------------------------------------ [2012-11-12 10:00:27] aharvey@php.net Automatic comment from SVN on behalf of aharvey Revision: http://svn.php.net/viewvc/?view=revision&revision=328321 Log: Update the Unicode character properties to document the existence of script names in PCRE. Fixes doc bug #54154 (Missing info about using \p{xx} and \P{xx} escape sequences with script names). ------------------------------------------------------------------------ [2011-03-03 22:27:56] hayk@php.net Description: ------------ Please update documentation and add information about using \p{xx} and \P{xx} escape sequences with script names. From http://www.pcre.org/pcre.txt When PCRE is built with Unicode character property support, three addi- tional escape sequences that match characters with specific properties are available. When not in UTF-8 mode, these sequences are of course limited to testing characters whose codepoints are less than 256, but they do work in this mode. The extra escape sequences are: \p{xx} a character with the xx property \P{xx} a character without the xx property \X an extended Unicode sequence The property names represented by xx above are limited to the Unicode script names, the general category properties, "Any", which matches any character (including newline), and some special PCRE properties (described in the next section). Other Perl properties such as "InMu- sicalSymbols" are not currently supported by PCRE. Note that \P{Any} does not match any characters, so always causes a match failure. Sets of Unicode characters are defined as belonging to certain scripts. A character from one of these sets can be matched using a script name. For example: \p{Greek} \P{Han} Those that are not part of an identified script are lumped together as "Common". The current list of scripts is: Arabic, Armenian, Avestan, Balinese, Bamum, Bengali, Bopomofo, Braille, Buginese, Buhid, Canadian_Aboriginal, Carian, Cham, Cherokee, Common, Coptic, Cuneiform, Cypriot, Cyrillic, Deseret, Devanagari, Egyp- tian_Hieroglyphs, Ethiopic, Georgian, Glagolitic, Gothic, Greek, Gujarati, Gurmukhi, Han, Hangul, Hanunoo, Hebrew, Hiragana, Impe- rial_Aramaic, Inherited, Inscriptional_Pahlavi, Inscriptional_Parthian, Javanese, Kaithi, Kannada, Katakana, Kayah_Li, Kharoshthi, Khmer, Lao, Latin, Lepcha, Limbu, Linear_B, Lisu, Lycian, Lydian, Malayalam, Meetei_Mayek, Mongolian, Myanmar, New_Tai_Lue, Nko, Ogham, Old_Italic, Old_Persian, Old_South_Arabian, Old_Turkic, Ol_Chiki, Oriya, Osmanya, Phags_Pa, Phoenician, Rejang, Runic, Samaritan, Saurashtra, Shavian, Sinhala, Sundanese, Syloti_Nagri, Syriac, Tagalog, Tagbanwa, Tai_Le, Tai_Tham, Tai_Viet, Tamil, Telugu, Thaana, Thai, Tibetan, Tifinagh, Ugaritic, Vai, Yi. Each character has exactly one Unicode general category property, spec- ified by a two-letter abbreviation. For compatibility with Perl, nega- tion can be specified by including a circumflex between the opening brace and the property name. For example, \p{^Lu} is the same as \P{Lu}. If only one letter is specified with \p or \P, it includes all the gen- eral category properties that start with that letter. In this case, in the absence of negation, the curly brackets in the escape sequence are optional; these two examples have the same effect: \p{L} \pL ------------------------------------------------------------------------ -- Edit this bug report at https://bugs.php.net/bug.php?id=54154&edit=1

« previous php.doc.bugs (#9097) next »