Bug #60616 [Csd]: odbc_fetch_into returns junk data at end of multi-byte char fields

From: Date: Mon, 28 Jul 2014 23:30:26 +0000
Subject: Bug #60616 [Csd]: odbc_fetch_into returns junk data at end of multi-byte char fields
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-186849@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=60616&edit=1 ID: 60616 Updated by: keyur@php.net Reported by: j dot faithw at yahoo dot com Summary: odbc_fetch_into returns junk data at end of multi-byte char fields Status: Closed Type: Bug Package: ODBC related Operating System: Linux PHP Version: 5.3.8 Assigned To: keyur Block user comment: N Private report: N New Comment: The fix for this bug has been committed. Snapshots of the sources are packaged every three hours; this change will be in the next snapshot. You can grab the snapshot at http://snaps.php.net/. For Windows: http://windows.php.net/snapshots/ Thank you for the report, and for helping us make PHP better. Previous Comments: ------------------------------------------------------------------------ [2014-07-28 23:29:06] keyur@php.net Thank you for your bug report. This issue has already been fixed in the latest released version of PHP, which you can download at http://www.php.net/downloads.php Patch committed to Git https://github.com/php/php-src/commit/00546bc9b7cb8208c27de9c6ebd12156401a5607 ------------------------------------------------------------------------ [2011-12-28 11:57:14] j dot faithw at yahoo dot com Description: ------------ This relates to bug#25792, which has been marked Analyzed but does not seem to be fixed. When retrieving data from a char() field containing multi-byte characters using e.g. odbc_fetch_into(), if the number of bytes used exceeds the number of characters then junk data is returned at the end of the string. I have experience this with Postgres char columns when the database is created with e.g the EUC_CN character encoding(createdb -E EUC_CN). This encoding uses between 1 and 3 bytes per character. So a char(10) could need up to 30 bytes. The problem is in the odbc_bindcols function in ext/odbc/php_odbc.c SQLColAttributes is called with SQL_COLUMN_DISPLAY_SIZE but this indicates the maximum number of characters required not the number of bytes. This means the buffer allocated for the value may not be big enough result->values[i].value=(char)emalloc(displaysize+1); Later on in e.g. odbc_fetch_into Z_STRLEN_P(tmp) = result->values[i].vallen; Z_STRVAL_P(tmp) = estrndup(result->values[i].value,Z_STRLEN_P(tmp)); This can result in a vallen bigger that displaysize. But the ODBC driver will only fill in at most displaysize+1 bytes(including null terminator). This means character data is missed and junk bytes are returned instead. The same problem may exist in ext/pdo_odbc/odbc_stmt.c. Where rc = SQLColAttribute(S->stmt, colno+1, SQL_DESC_DISPLAY_SIZE, NULL, 0, NULL, &displaysize); is called. But I have not tested this. The following fixes odbc_bindcols for the char(x) datatype. I believe 4 bytes is the maximum required for any character encoding. php_odbc.c:line 988 if (result->values[i].coltype == SQL_CHAR) { //If using a multibyte character encoding //number of bytes could be 4*SQL_COLUMN_DISPLAY_SIZE. //Without this workaround various functions //e.g. odbc_fetch_into will return data with a null after //diplaysize bytes and extra junk data at the end as //vallen can be bigger than displaysize. Tested using //PostgreSQL with EUC_CN encoding. displaysize*=4; } The fix may be needed for other data types as well as SQL_CHAR. ------------------------------------------------------------------------ -- Edit this bug report at https://bugs.php.net/bug.php?id=60616&edit=1

« previous php.bugs (#186849) next »