Re: Unicode strings?
| From: | Crypto Compress | Date: | Wed, 12 Mar 2014 09:49:57 +0000 |
| Subject: | Re: Unicode strings? | ||
| References: | 1 2 3 | Groups: | php.internals |
| Request: | Send a blank email to internals+get-73073@lists.php.net to get a copy of this message | ||
Hi,
Am 11.03.2014 13:27, schrieb Lester Caine:
Crypto Compress wrote:Quote #1: "You can request 64 or 32 bits with the --with-library-bits= option, ..." Quote #2: "Strings are represented as UChar * as the base string type." http://userguide.icu-project.org/icufaq#TOC-How-do-I-get-32--or-64-bit-versions-of-the-ICU-libraries- String length is platform dependent.This information has been published in several places on the list and in the wiki already ... http://userguide.icu-project.org/strings/utf-8 for the ICU, and the RFC's here for 64 bit improvements to PHP ...I'm slowly working through a long list of things relating to unicode strings trying to work out just where the main problems are. The very first problem I hit is ICU's limitation to 32bit string lengths. How does the switch to 64bit string length on 64 bit platforms impinge on this. While I can see the advantage of this particular change, would that also now require our own version of ICU capable of also handling longer strings? This probably falls out in the wash of my next point ...Where have you found this information? Can you please provide source for this?
I think of this as a "immutable ValueObject". If a string is converted, there is no reason to cache the old string. binary => convert to utf-8 as de_de.iso-8859-15@euro => {"utf-8", "de_DE_EURO", binary} => convert to utf-32 => {"utf-32", "de_DE_EURO", bigger-binary} What other data is needed in here to be doubtless unicode? Is locale needed at all? Should it be nullable? Case-(in)sensitive? cryptocompressHow do you provide a holder for the various additional items required for a unicode 'object'? While I can see one would get away with calling functions all the time on a single string object, having calculated different versions of the same string or complex character counts, they need to be cached so they can be used again? Or does one maintain each answer in different variables?Currently strings are simply strings? I'm sure we have already had this discussion, and it will be necessary to switch from simple strings to a string object which can handle the intricacies of unicode?Yes, currently we have so called binary strings (simple bytes, 8 bits). No, we should not create an string-object to handle all intricacies of unicode.