Req #73716 [Opn->Fbk]: PHP 7.1's CHCP switching in console should be optional

From: Date: Tue, 13 Dec 2016 17:41:54 +0000
Subject: Req #73716 [Opn->Fbk]: PHP 7.1's CHCP switching in console should be optional
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-205964@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=73716&edit=1 ID: 73716 Updated by: ab@php.net Reported by: anrdaemon at freemail dot ru Summary: PHP 7.1's CHCP switching in console should be optional -Status: Open +Status: Feedback Type: Feature/Change Request Package: Output Control Operating System: Windows PHP Version: 7.1.0 Block user comment: N Private report: N New Comment: @anrdaemon, it would be nice, if you could put your language down to the technical level from your shame theory. You say " half the extensions can't be loaded" - could you post some reproduce case for this, that can be debugged? Thanks. Previous Comments: ------------------------------------------------------------------------ [2016-12-13 16:01:03] anrdaemon at freemail dot ru That's just reinforcing my point, that single-minded pseudosolutions are not going to cut it. Terminal encoding is a known complex problem, and by now people developed a treasure trove of knowledge in dealing with it. Why PHP has to "invent" its own ways? The "one size fits all" approach didn't quite worked for several last centuries only on my memory. I don't want to see anyone trying to walk that road only to be covered in shame. Yet again. ------------------------------------------------------------------------ [2016-12-13 07:14:25] yohgaki@php.net Since you seems to using ISO-8859 compatible encoding, your situation is better than CP932(SJIS). You can simply use CP866, but CP932 cannot be internal_encoding. Therefore, we(Japanese) has to use UTF-8/EUC-JP for internal_encoding on Windows. Although, web input/output conversion can be handled by mbstring/iconv, inputs/outputs have to converted manually, filenames especially. For this reason, I suppose most Japanese PHP users do not use multibyte filenames with PHP. If you could use your code page without problems, I suggest to use it as internal_encoding with CLI. It's a lot easier. You may try cmd.exe /f:on /k "chcp 65001 to use UTF-8 with cmd.exe. I tried it on my Windows. Japanese filenames (CP932) raised error and stopped "dir"ing. If all of your filenames are UTF-8, it may work. However, other programs like explorer may have problems with UTF-8 filenames. (I don't know how your version of Windows behave) Anyway, feasible resolution for mixed encoding environment that treats various encoding automagically is very tough subject and use of UTF-8 could be problematic as described above. ------------------------------------------------------------------------ [2016-12-13 05:42:28] anrdaemon at freemail dot ru Said that, let's take *NIX as example. When you are writing to terminal in *NIX, you don't suddenly change terminal codepage, you translate your data from your program's internal codepage to the terminal's one. Why on earth Windows terminal has to be any different? I can imagine the time it took for you to write all that text, but it's senseless. How's "UTF-8 is undesirable"? It IS desirable. Internally. I want my application to use UTF-8 wherever possible. Emphasis on "possible". As opposed to "wherever it want regardless of my expectations". I'm reading your argumentations and all my reaction is an urge to shake my head in an attempt to get the wrongs out of it. ------------------------------------------------------------------------ [2016-12-13 05:04:08] anrdaemon at freemail dot ru World doesn't rotate around PHP, and there's other programs writing to the same console at the same time. Which totally do not expect random CP changes. Least of all, I totally do not expect CP changes from INTERNAL program settings. ------------------------------------------------------------------------ [2016-12-12 21:30:39] ab@php.net A bit more background regarding these behaviors. The world is UTF-8 today. The Windows path issue, both long and UTF-8, was long standing. UTF-8 is now default in PHP on Windows, like it became relatively long ago on Linux and other platforms, and in PHP-5.6 not very long ago. Still Windows is very different from other internationalization approaches, despite some improvements in system locale handling are to see. Many APIs are still codepage bound, including console, path, I/O, etc. That is unlikely to change soon, if at all. This makes the portability of PHP on Windows itself to lag. So for one - the behavior in PHP needs to be more consolidated across platform, for the other - it needs to be done a simple way. But even then - there are various platform issues, so then the actual thing is sometimes tricky to stretch straight. @anrdiemon, the INI configuration listed is inconsistent for 7.1 and even for earlier. For 7.1, it diverges from what was documented in UPGRADING in first place. Then, the output_encoding directive is only useful, if the usage of the iconv/mbstring ob handler use is intended. As Yasuo mentioned, this ob handler has no effect on CLI. Furthermore - any of *_encoding are deprecated, see http://php.net/manual/en/iconv.configuration.php Here's what i have with a non existent extension DLL [code] C:\php-sdk\php71\vc14\x64\php-src $ chcp Active code page: 437 C:\php-sdk\php71\vc14\x64\php-src $ x64\Release\php.exe -n -d extension_dir=nowhere -d extension=notfound -v PHP Warning: PHP Startup: Unable to load dynamic library 'nowhere\notfound' - The specified module could not be found. in Unknown on line 0 PHP 7.1.1-dev (cli) (built: Dec 12 2016 15:15:56) ( NTS MSVC14 (Visual C++ 2015) x64 ) Copyright (c) 1997-2016 The PHP Group Zend Engine v3.1.0, Copyright (c) 1998-2016 Zend Technologies C:\php-sdk\php71\vc14\x64\php-src $ chcp Active code page: 437 [/code] The encoding determination sequence, as specifically documented in UPGRADING, is kept same, despite internal_encoding is used. Otherwise, there is no behavior difference, neither in earlier PHP version, nor on another platform. What is new - yes, the console codepage is switched automatically, but that is not without a reason. The console is UTF-8 on the overwhelming number of platforms. @requinix, that's a good catch. Of course, if a process is sent a KILL, it won't be able to handle it. Same will happen when SIGSEGV and several other situations occur. This kind of behavior is what i expect @anrdaemon experiences. Any controlled exit from will sure restore the console, but if -9 is sent - there's nothing that can be done. So this is a real crash in a C program, that have to be happening. As for me, adding an INI to just workaround a crash, is not sensible. And probably, doing this colud be even misleading. Imagine, you output a cyrillic string with 7.1, while the console codepage is 437 - [code] C:\php-sdk\php71\vc14\x64\php-src $ chcp Active code page: 437 C:\php-sdk\php71\vc14\x64\php-src $ x64\Release\php.exe -n -d default_charset="CP866" -r "var_dump(sapi_windows_cp_get(), 'привет');" int(866) string(6) "привет" C:\php-sdk\php71\vc14\x64\php-src $ chcp Active code page: 437 C:\php-sdk\php71\vc14\x64\php-src $ [/code] If there were default_charset=cp437 directive set to 7.1, what it gave is this - string(6) "??????", just like 7.0 would do by default. With default_charset=UTF-8 as default in 7.0, it were same, but in 7.1, it's [code] string(12) "привет" int(65001) [/code] Now, if one would want to out another one, say a Czech string - this is broken again. The only what works is - using UTF-8 console output. This conserns same for direct input, or for warnings, error messages, open and output paths, getting various data like user names, et cetera. Even there were an INI, and the output would be turned off, either one or the other of the cases will put mojibake onto console. Other side - the input will be in incompatible incoding. If the console is, say cp866, but PHP uses UTF-8 internaly, and you want to read a filename from console to put it into an I/O functionin. PHP will internally try to convert the cp866 char's into wchar_t's using utf-8. Now, we can of course say - lets use different codepage for input, than convert it to internal, then convert again into output. Diverging codepages for PHP internal, PHP input, PHP output, console input, console output ... well, hopefully one can see where it leads. Subsequent weirdness and bug reports! are guaranteed :) The way to do it, if UTF-8 is not desired, is simply setting like internal_encoding=cp1251, which will keep the current console codepage but also disable the multibyte path and other feature support. In this case - all the behavior is turned to what it was before 7.1. This is documented and done this way to explicitly keep the the backward compatibility for older apps or for scripts requiring the old behavior. Sure, any implementation can't be perfect enough, this one exhausts the most of the possibilities systems provide. If interested, it were also worth it to check, how this topic is handled in other language, Python for example :) @anrdaemon, the reported issue is something, that is caused by not following the recommended UPGRADING way. The ini configuration is not supposed to deliver the expected result in any case. With the crash behavior as described by @requinix- yeah, that's a known one, but it is something different. It is noticeable in an abnormal crash situation and is impossible to catch. Any other program won't behave different in this case. Now, if we say, PHP crashes that often, that it becomes an issue of this kind - then we have something else to fix. UTF-8 and the wide APIs usage is what really matters for the future and is the worthy goal to strive. The old PHP on Windows behavior is still available by putting the corresponding configuration. UTF-8 became default in 5.6, now it's reflected on the Windows side as well. There is a number of factors, that can make UTF-8 usage not as easy as on other platforms. There are Windows specific things, and there is some learn curve, but there is no reason to mix the old and new behavior. Either an app has a clean UTF-8 support, or it goes by the legacy behavior. A "half UTF-8" support an over complicated implementation would be something weird, as for me. If there's a crash scenario to investigate, that'd what should be done. But otherwise, i'd see it as "not a bug", same as Yasuo. Thanks. ------------------------------------------------------------------------ The remainder of the comments for this report are too long. To view the rest of the comments, please view the bug report online at https://bugs.php.net/bug.php?id=73716 -- Edit this bug report at https://bugs.php.net/bug.php?id=73716&edit=1

« previous php.bugs (#205964) next »