Bug #60616 [Csd]: odbc_fetch_into returns junk data at end of multi-byte char fields
| From: | keyur@php.net | Date: | Mon, 28 Jul 2014 23:30:26 +0000 |
| Subject: | Bug #60616 [Csd]: odbc_fetch_into returns junk data at end of multi-byte char fields | ||
| References: | 1 | Groups: | php.bugs |
| Request: | Send a blank email to php-bugs+get-186849@lists.php.net to get a copy of this message | ||
Edit report at https://bugs.php.net/bug.php?id=60616&edit=1
ID: 60616
Updated by: keyur@php.net
Reported by: j dot faithw at yahoo dot com
Summary: odbc_fetch_into returns junk data at end of
multi-byte char fields
Status: Closed
Type: Bug
Package: ODBC related
Operating System: Linux
PHP Version: 5.3.8
Assigned To: keyur
Block user comment: N
Private report: N
New Comment:
The fix for this bug has been committed.
Snapshots of the sources are packaged every three hours; this change
will be in the next snapshot. You can grab the snapshot at
http://snaps.php.net/.
For Windows:
http://windows.php.net/snapshots/
Thank you for the report, and for helping us make PHP better.
Previous Comments:
------------------------------------------------------------------------
[2014-07-28 23:29:06] keyur@php.net
Thank you for your bug report. This issue has already been fixed
in the latest released version of PHP, which you can download at
http://www.php.net/downloads.php
Patch committed to Git https://github.com/php/php-src/commit/00546bc9b7cb8208c27de9c6ebd12156401a5607
------------------------------------------------------------------------
[2011-12-28 11:57:14] j dot faithw at yahoo dot com
Description:
------------
This relates to bug#25792, which has been marked Analyzed but does not seem to be fixed.
When retrieving data from a char() field containing multi-byte characters using e.g.
odbc_fetch_into(), if the number of bytes used exceeds the number of characters then junk data is
returned at the end of the string.
I have experience this with Postgres char columns when the database is created with e.g the EUC_CN
character encoding(createdb -E EUC_CN).
This encoding uses between 1 and 3 bytes per character. So a char(10) could need up to 30 bytes.
The problem is in the odbc_bindcols function in ext/odbc/php_odbc.c
SQLColAttributes is called with SQL_COLUMN_DISPLAY_SIZE but this indicates the maximum number of
characters required not the number of bytes.
This means the buffer allocated for the value may not be big enough
result->values[i].value=(char)emalloc(displaysize+1);
Later on in e.g. odbc_fetch_into
Z_STRLEN_P(tmp) = result->values[i].vallen;
Z_STRVAL_P(tmp) = estrndup(result->values[i].value,Z_STRLEN_P(tmp));
This can result in a vallen bigger that displaysize. But the ODBC driver will only fill in at most
displaysize+1 bytes(including null terminator). This means character data is missed and junk bytes
are returned instead.
The same problem may exist in ext/pdo_odbc/odbc_stmt.c. Where
rc = SQLColAttribute(S->stmt, colno+1, SQL_DESC_DISPLAY_SIZE,
NULL, 0, NULL, &displaysize);
is called. But I have not tested this.
The following fixes odbc_bindcols for the char(x) datatype. I believe 4 bytes is the maximum
required for any character encoding.
php_odbc.c:line 988
if (result->values[i].coltype == SQL_CHAR) {
//If using a multibyte character encoding
//number of bytes could be 4*SQL_COLUMN_DISPLAY_SIZE.
//Without this workaround various functions
//e.g. odbc_fetch_into will return data with a null after
//diplaysize bytes and extra junk data at the end as
//vallen can be bigger than displaysize. Tested using
//PostgreSQL with EUC_CN encoding.
displaysize*=4;
}
The fix may be needed for other data types as well as SQL_CHAR.
------------------------------------------------------------------------
--
Edit this bug report at https://bugs.php.net/bug.php?id=60616&edit=1