Re: [RFC] UString
| From: | Rowan Collins | Date: | Wed, 22 Oct 2014 08:09:50 +0000 |
| Subject: | Re: [RFC] UString | ||
| References: | 1 | Groups: | php.internals |
| Request: | Send a blank email to internals+get-78219@lists.php.net to get a copy of this message | ||
On 21 October 2014 23:21:37 GMT+01:00, Andrea Faulds <ajf@ajf.me> wrote:
>Make array-like indexing with [] be by
>code points as you may be able to do that in constant time
If the internal representation is UTF8, both code point and grapheme access require traversal unless
you have some additional index structure. Both can be trivialised to byte access if you have
detected and stored that the string is entirely ASCII, but otherwise you will nearly always have
multiple widths within one string.
If the internal representation is UTF16, code point access can be accelerated for any string
containing only BMP characters (no surrogate pairs). The Perl6 concept of "NFG" attempts
to extend that advantage to grapheme access, and to points outside the BMP.