Re: Re: strlen() under unicode.semantics
| From: | Andrei Zmievski | Date: | Fri, 23 Jun 2006 06:37:31 +0000 |
| Subject: | Re: Re: strlen() under unicode.semantics | ||
| References: | 1 | Groups: | php.internals |
| Request: | Send a blank email to internals+get-24191@lists.php.net to get a copy of this message | ||
The only way they can get at the internal UTF-16 representation is via unicode_encode($uni, 'UTF-16') which will return a binary UTF-16 string. In that case, strlen() will work just as well.
-Andrei
On Jun 22, 2006, at 11:30 PM, Andi Gutmans wrote:
I don't quite agree. I think there's a good chance people will want to save Unicode strings in a binary format for performance reasons. Save it the way it looks in memory, and put it back... Why convert to UTF-8 or any other encoding if it's just about storage? Andi-----Original Message----- From: Sara Golemon [mailto:pollita@php.net] Sent: Thursday, June 22, 2006 9:15 PM To: "Ron Korving" Cc: internals@lists.php.net Subject: Re: [PHP-DEV] Re: strlen() under unicode.semantics--PHP Internals - PHP Runtime Development Mailing List To unsubscribe, visit: http://www.php.net/unsub.phpStill, it's gotta be useful to be know how many bytes it occupies. Perhaps for Content-length headers or something. There areplenty oflow level concepts to think of where one might need this.And even ifyou can't think of any reason now, you don't wanna get hitin the faceby it and have to implement such a function for PHP 6.0.1.For this type of usage, I'd think it'd be relevant to know how many bytes the string will occupy in a given output encoding moreso that what it happens to occupy in the underlying implementation. In the example you cited, string contents will more typically be sent as utf8 rather than the utf16 of php's internal encoding. $utf8str = unicode_encode($unistr, 'utf8'); header('Content-type: text/html; encoding="utf8"'); header('Content-length: ' . strlen($utf8str)); echo $utf8str; I'm not saying it's impossible that a legitimate use will come up to know the internal byte-usage of a unicode string, there's certainly no harm in adding such a function (apart from the tired shot-foot argument). I just doubt you (or anyone) will come up such a case anytime soon. -Sara -- PHP Internals - PHP Runtime Development Mailing List To unsubscribe, visit: http://www.php.net/unsub.php