Doc #31740 [Com]: fgetcsv skips fields that start with an umlaut character

From: Date: Wed, 25 Sep 2013 15:19:03 +0000
Subject: Doc #31740 [Com]: fgetcsv skips fields that start with an umlaut character
References: 1  Groups: php.doc.bugs 
Request: Send a blank email to doc-bugs+get-10304@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=31740&edit=1 ID: 31740 Comment by: webspam at live dot de Reported by: arjan at avoid dot org Summary: fgetcsv skips fields that start with an umlaut character Status: Closed Type: Documentation Problem Package: Documentation problem Operating System: Linux (Suse) PHP Version: 5.0.3 Block user comment: N Private report: N New Comment: In PHP 5.3.2 this problem is still there. No solution? Previous Comments: ------------------------------------------------------------------------ [2011-01-30 15:29:13] max dot wildgrube at web dot de Again: the 1st-umlaut-vasnishs problem: (PHP version 5.2.6) Sorry, but for many of non-English web developers this "solution" is not helpful. As for those of us who are not a system administrator has the problem that we cannot influence the settings of PHP or Apache. So in my environment the "Safe mode" was switched on, preventing the usage of putenv. And I think it is an illusion to dispose some of the big providers to switch the safe mode off (even if this feature is DEPRECATED). And $_ENV ['LANG'] = 'en_US' does not healed the problem (nor setting de_DE). Nevertheless the environment variable LANG is not set (asking getenv, $_ENV, phpinfo). And I think this problem has nothing to do with the encoding: Other “inline” umlauts are preserved as estimated. If the data field is enclosed with quotes the 1st Umlaut after the introducing quote (therefore the 2nd character) survives. So I live now with the ugly workaround to place a magic sequence ~~ before every 1st umlaut in the csv file with: preg_replace ("/(\t)([€-ÿ])/", "\t~~$2", $import); and remove these sequences after fgetcsv while reading the field array with: foreach ( $columns as $col ) { $col = trim ($col, '~~'); … Max. ------------------------------------------------------------------------ [2005-02-04 11:26:03] vrana@php.net This bug has been fixed in the documentation's XML sources. Since the online and downloadable versions of the documentation need some time to get updated, we would like to ask you to be a bit patient. Thank you for the report, and for helping us make our documentation better. "Locale setting is taken into account by this function. If LANG is e.g. en_US.UTF-8, files in one-byte encoding are read wrong by this function." ------------------------------------------------------------------------ [2005-01-29 23:22:22] derick@php.net Yeah, we should have some information that tells people that the locale setting have effect on this, recategorizing... ------------------------------------------------------------------------ [2005-01-29 12:07:35] arjan at avoid dot org The LANG environment variable on the faulty machines was set like this: LANG="en_US.UTF-8" When changed to LANG="en_US" The problem is fixed. Thanks a lot! However, shouldn't this behaviour be mentioned in the manual for fgetcsv? I can imagine more people experiencing this 'bug' that turns out to be not a bug... Thanks again! ------------------------------------------------------------------------ [2005-01-29 03:33:57] moriyoshi@php.net What locale specifier is set to LANG or LC_CTYPE environment variable? ------------------------------------------------------------------------ The remainder of the comments for this report are too long. To view the rest of the comments, please view the bug report online at https://bugs.php.net/bug.php?id=31740 -- Edit this bug report at https://bugs.php.net/bug.php?id=31740&edit=1

« previous php.doc.bugs (#10304) next »