Re: Re: strlen() under unicode.semantics
| From: | Andrei Zmievski | Date: | Fri, 23 Jun 2006 07:20:56 +0000 |
| Subject: | Re: Re: strlen() under unicode.semantics | ||
| References: | 1 | Groups: | php.internals |
| Request: | Send a blank email to internals+get-24196@lists.php.net to get a copy of this message | ||
Really? I think it's very rare that someone'd want to get at the internals of a Unicode string.
-Andrei
On Jun 22, 2006, at 11:44 PM, Andi Gutmans wrote:
Hmm, I was thinking we might have some binary write function which would do that automagically. I think it'd be worth it.-----Original Message----- From: Andrei Zmievski [mailto:andrei@gravitonic.com] Sent: Thursday, June 22, 2006 11:38 PM To: Andi Gutmans Cc: 'Sara Golemon'; '"Ron Korving"'; internals@lists.php.net Subject: Re: [PHP-DEV] Re: strlen() under unicode.semantics The only way they can get at the internal UTF-16 representation is via unicode_encode($uni, 'UTF-16') which will return a binary UTF-16 string. In that case, strlen() will work just as well. -Andrei On Jun 22, 2006, at 11:30 PM, Andi Gutmans wrote:I don't quite agree. I think there's a good chance peoplewill want tosave Unicode strings in a binary format for performancereasons. Saveit the way it looks in memory, and put it back... Whyconvert to UTF-8or any other encoding if it's just about storage? Andihow many-----Original Message----- From: Sara Golemon [mailto:pollita@php.net] Sent: Thursday, June 22, 2006 9:15 PM To: "Ron Korving" Cc: internals@lists.php.net Subject: Re: [PHP-DEV] Re: strlen() under unicode.semanticsStill, it's gotta be useful to be know how many bytes it occupies. Perhaps for Content-length headers or something. There areplenty oflow level concepts to think of where one might need this.And even ifyou can't think of any reason now, you don't wanna get hitin the faceby it and have to implement such a function for PHP 6.0.1.For this type of usage, I'd think it'd be relevant to knowmoreso thatbytes the string will occupy in a given output encodingimplementation. In thewhat it happens to occupy in the underlyingcome up toexample you cited, string contents will more typically be sent as utf8 rather than the utf16 of php's internal encoding. $utf8str = unicode_encode($unistr, 'utf8'); header('Content-type: text/html; encoding="utf8"'); header('Content-length: ' . strlen($utf8str)); echo $utf8str; I'm not saying it's impossible that a legitimate use willcertainlyknow the internal byte-usage of a unicode string, there'sunsubscribe,no harm in adding such a function (apart from the tired shot-foot argument). I just doubt you (or anyone) will come up such a case anytime soon. -Sara -- PHP Internals - PHP Runtime Development Mailing List Tounsubscribe,visit: http://www.php.net/unsub.php-- PHP Internals - PHP Runtime Development Mailing List Tovisit: http://www.php.net/unsub.php