Bug #18169 Updated: Driver cannot deliver UCS-2 unicode to SQL Server

From: Date: Fri, 05 Jul 2002 08:21:52 +0000
Subject: Bug #18169 Updated: Driver cannot deliver UCS-2 unicode to SQL Server
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-13189@lists.php.net to get a copy of this message
 ID:               18169
 Updated by:       joesterg@hotmail.com
 Reported By:      joesterg@hotmail.com
 Status:           Open
 Bug Type:         MSSQL related
 Operating System: Windows 2000 Server
 PHP Version:      4.1.2
 New Comment:

You are probably right. However, Unicode is central to making
world-wide web applications, and all major database vendors have this
posibility.
I find it to be a hindrance to wider deployment of large-scale,
worldwide php applications.

Does anyone know if it is only the MSSQL module? -are there any plans
to look into this issue?

What are the future directions for PHP and Unicode support?


Previous Comments:
------------------------------------------------------------------------

[2002-07-05 04:14:38] yohgaki@php.net

Wide char encoding, UCS2/UCS4/UTF16/UTF32, don't work well with current
PHP. I guess SQL Server module is using strlen() or like, that cannot
be used with wide char...

Fixing this is not simple at all. 


------------------------------------------------------------------------

[2002-07-04 18:10:24] joesterg@hotmail.com

I have a problem converting UTF-8 (web character encoding) to UCS2
(Microsoft Windows character encoding) using PHP, and storing this in
the Microsoft SQL Server 2000 database.

My setup is:
Windows 2000 Server, with Apache 1.3.24/PHP 4.1.1 and Microsoft SQL
Server 2000

Now, as a result of Microsofts Q232580, I will have to do conversion
between UTF-8 and UCS-2. For this, I thought I would use the Multibyte
String functions.
However, this does not seem to work.

I am absolutely sure, that I input UTF-8 encoded data into my string,
and then I do:
$ucs2string=mb_convert_encoding($string,"UCS2","UTF-8");
$sqlStmt="insert into testtbl (tekst) values(N'".($ucs2string)."')";
$rs=$DBCon->Execute($sqlStmt);

When I access the database, then I will see something stored, that does
not resemble the input at all (most times, I see Japanese/Chinese
characters?!??). Furthermore, the insert sometimes comes up with an
error, and consequently stores nothing.

To me, it seems like either one of these (or both) are flawed:
1. the Multibyte String encoding funtion does not work properly (ie.
encoding from UTF-8 to UCS-2 does not happen correctly).
2. The PHP MSSQL driver does not handle unicode data properly, even
though the target column in the database is specified as Unicode and N
is prepended to the string before insert.

This leads me to use ADO (as in the example above), storing UTF-8
encoded data into SQL Server -this is a very short term solution, as
data are not sortable in the database (some of it looks like garbage
because of the
missing encoding).


------------------------------------------------------------------------


-- 
Edit this bug report at http://bugs.php.net/?id=18169&edit=1



Thread (14 messages)

« previous php.bugs (#13189) next »