Edit report at http://bugs.php.net/bug.php?id=31740&edit=1
ID: 31740
Comment by: max dot wildgrube at web dot de
Reported by: arjan at avoid dot org
Summary: fgetcsv skips fields that start with an umlaut
character
Status: Closed
Type: Documentation Problem
Package: Documentation problem
Operating System: Linux (Suse)
PHP Version: 5.0.3
Block user comment: N
Private report: N
New Comment:
Again: the 1st-umlaut-vasnishs problem:
(PHP version 5.2.6)
Sorry, but for many of non-English web developers this "solution" is not
helpful. As for those of us who are not a system administrator has the
problem that we cannot influence the settings of PHP or Apache. So in my
environment the "Safe mode" was switched on, preventing the usage of
putenv. And I think it is an illusion to dispose some of the big
providers to switch the safe mode off (even if this feature is
DEPRECATED).
And $_ENV ['LANG'] = 'en_US' does not healed the problem (nor setting
de_DE).
Nevertheless the environment variable LANG is not set (asking getenv,
$_ENV, phpinfo).
And I think this problem has nothing to do with the encoding: Other
âinlineâ umlauts are preserved as estimated.
If the data field is enclosed with quotes the 1st Umlaut after the
introducing quote (therefore the 2nd character) survives.
So I live now with the ugly workaround to place a magic sequence ~~
before every 1st umlaut in the csv file with:
preg_replace ("/(\t)([â¬-ÿ])/", "\t~~$2", $import);
and remove these sequences after fgetcsv while reading the field array
with:
foreach ( $columns as $col ) { $col = trim ($col, '~~'); â¦
Max.
Previous Comments:
------------------------------------------------------------------------
[2005-02-04 11:26:03] vrana@php.net
This bug has been fixed in the documentation's XML sources. Since the
online and downloadable versions of the documentation need some time
to get updated, we would like to ask you to be a bit patient.
Thank you for the report, and for helping us make our documentation
better.
"Locale setting is taken into account by this function. If LANG is e.g.
en_US.UTF-8, files in one-byte encoding are read wrong by this
function."
------------------------------------------------------------------------
[2005-01-29 23:22:22] derick@php.net
Yeah, we should have some information that tells people that the locale
setting have effect on this, recategorizing...
------------------------------------------------------------------------
[2005-01-29 12:07:35] arjan at avoid dot org
The LANG environment variable on the faulty machines was set like this:
LANG="en_US.UTF-8"
When changed to
LANG="en_US"
The problem is fixed. Thanks a lot!
However, shouldn't this behaviour be mentioned in the
manual for fgetcsv? I can imagine more people experiencing
this 'bug' that turns out to be not a bug...
Thanks again!
------------------------------------------------------------------------
[2005-01-29 03:33:57] moriyoshi@php.net
What locale specifier is set to LANG or LC_CTYPE
environment variable?
------------------------------------------------------------------------
[2005-01-28 20:28:51] arjan at avoid dot org
In order to narrow the problem down as much as I can, I tried the
following script as well on the system that have problems with fgetcsv:
<?php
$fp = fopen('csv_test.csv', 'r');
while (!feof($fp)) {
$buffer = fgets($fp, 4096);
echo $buffer;
}
fclose($fp);
?>
In this case, the umlauts do get read and printed.
------------------------------------------------------------------------
The remainder of the comments for this report are too long. To view
the rest of the comments, please view the bug report online at
http://bugs.php.net/bug.php?id=31740--
Edit this bug report at http://bugs.php.net/bug.php?id=31740&edit=1