Bug #66496 [Fbk->Dup]: conversion of UTF-8 strings containing language specific chars is wrong

From: Date: Tue, 29 Apr 2014 11:51:46 +0000
Subject: Bug #66496 [Fbk->Dup]: conversion of UTF-8 strings containing language specific chars is wrong
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-185499@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=66496&edit=1

 ID:                 66496
 Updated by:         ab@php.net
 Reported by:        care at novadys dot de
 Summary:            conversion of UTF-8 strings containing language
                     specific chars is wrong
-Status:             Feedback
+Status:             Duplicate
 Type:               Bug
 Package:            COM related
 Operating System:   Windows
 PHP Version:        5.5.8
 Block user comment: N
 Private report:     N

 New Comment:

This is fixed now as it's the same as bug #66431, please check.

Thanks.


Previous Comments:
------------------------------------------------------------------------
[2014-01-17 19:52:14] ab@php.net

We need a code snippet in PHP to fix and write a test.

Thanks.

------------------------------------------------------------------------
[2014-01-17 14:56:54] care at novadys dot de

It works correctly when replacing the line

V_BSTR(v) = SysAllocStringByteLen((char*)olestring, Z_STRLEN_P(z) * sizeof(OLECHAR));

by

V_BSTR(v) = SysAllocStringByteLen((char*)olestring, wcslen(olestring) * sizeof(OLECHAR));

------------------------------------------------------------------------
[2014-01-16 20:28:26] ab@php.net

Were it possible you to post a repro snippet? Thanks.

------------------------------------------------------------------------
[2014-01-16 14:43:19] care at novadys dot de

Description:
------------
Using a UTF-8 String as input for a com-function will generate a wrong string in COM interface. If
the input string is containing n-chars which are encoded with 2 bytes, the length of resulting
string is n byte too long. 

example: "I want to Dusseldorf and Koln" is correctly handled
"I want to Düsseldorf and Köln" will call COM function with a string:
"I want to Düsseldorf and Köln\0\4"  

reason:

ext/com_dotnet/com_variant.c

...
PHP_COM_DOTNET_API void php_com_variant_from_zval(VARIANT *v, zval *z, int codepage TSRMLS_DC)

...

case IS_STRING:
                        V_VT(v) = VT_BSTR;
                        olestring = php_com_string_to_olestring(Z_STRVAL_P(z), Z_STRLEN_P(z),
codepage TSRMLS_CC);
                        
here is the problem:

V_BSTR(v) = SysAllocStringByteLen((char*)olestring, Z_STRLEN_P(z) * sizeof(OLECHAR));

When input string is UTF-8 encoded Z_STRLEN_P(z) has a count of 2 byte for each "special"
char. So length of input string is count of all chars + count of "special" chars. After
conversion to olestring Z_STRLEN_P(z) is the wrong length the olestring is shorter. So
SysAllocStringByteLen is allocating too much memory and while the for the string length of
olestring. In result those strings are always showing a '\0' for the first special char +
n-1 random chars for the remaing special chars.






------------------------------------------------------------------------



-- 
Edit this bug report at https://bugs.php.net/bug.php?id=66496&edit=1


Thread (5 messages)

« previous php.bugs (#185499) next »