[php-src] Issue #8281: mb_convert_encoding "\" (backslash) and "~" (tilde) convert failed to Shift_JIS

From: Date: Tue, 05 Apr 2022 00:31:23 +0000
Subject: [php-src] Issue #8281: mb_convert_encoding "\" (backslash) and "~" (tilde) convert failed to Shift_JIS
Groups: php.bugs 
Request: Send a blank email to php-bugs+get-240660@lists.php.net to get a copy of this message
Issue: https://github.com/php/php-src/issues/8281 Comment Author: youkidearitai There are tons of cases where Unicode and Shift_JIS conversions are involved in CSV uploads and downloads, and it can be difficult to find out how much they are. Similarly, it's hard to find out how much a Windows path doesn't work. However, many users will find it very difficult for 0x5C and 0x7E to have "strict" conversions the moment they upgrade to PHP 8.1. This is good enough for Japanese users to hesitate to upgrade to PHP 8.1. This is because it is perceived by the Japanese as being converted to a different character. > It is certainly confusing. My preference is to follow published specifications when possible; I > believe this tends to reduce confusion in the long term. However, if there are strong practical > reasons to deviate from specifications, that can certainly be done. > However, we do not want to flip-flop back and forth. To avoid flip-flopping, we need to > thoroughly understand all the implications of either following the spec or deviating from it. After > all factors are considered, and as many interested parties as possible are consulted, if the final > decision is to change, then we should document the reason for the decision and stick to it. As you said, I think it is correct to follow the published specifications. However, this is a change that breaks backwards compatibility, and if so, I feel that this change should be discussed in PHP RFCs and so on. As about for the convention(customary), at least in Python 3, even if 0x5C or 0x7E is converted to Shift_JIS, it is converted as it is. $ python3 Python 3.8.10 (default, Mar 15 2022, 12:22:08) [GCC 9.4.0] on linux Type "help", "copyright", "credits" or "license" for more information. >>> '\\'.encode("SJIS") b'\\' >>> '\\~'.encode("SJIS") b'\\~' >>> 'あ'.encode("SJIS") b'\x82\xa0' >>> I tried it with Ruby 2.7. After all it converted as it is. $ ruby --version ruby 2.7.0p0 (2019-12-25 revision 647ee6f091) [x86_64-linux-gnu] $ ruby -e 'puts "あ".encode("SJIS").encode("UTF-8")' あ $ ruby -e 'puts "\\".encode("SJIS")' \ $ ruby -e 'puts "~".encode("SJIS")' ~

« previous php.bugs (#240660) next »