[php-src] Issue #8281: mb_convert_encoding "\" (backslash) and "~" (tilde) convert failed to Shift_JIS
| From: | youkidearitai | Date: | Tue, 05 Apr 2022 00:31:23 +0000 |
| Subject: | [php-src] Issue #8281: mb_convert_encoding "\" (backslash) and "~" (tilde) convert failed to Shift_JIS | ||
| Groups: | php.bugs | ||
| Request: | Send a blank email to php-bugs+get-240660@lists.php.net to get a copy of this message | ||
Issue: https://github.com/php/php-src/issues/8281
Comment Author: youkidearitai
There are tons of cases where Unicode and Shift_JIS conversions are involved in CSV uploads and
downloads, and it can be difficult to find out how much they are.
Similarly, it's hard to find out how much a Windows path doesn't work.
However, many users will find it very difficult for 0x5C and 0x7E to have "strict"
conversions the moment they upgrade to PHP 8.1. This is good enough for Japanese users to hesitate
to upgrade to PHP 8.1.
This is because it is perceived by the Japanese as being converted to a different character.
> It is certainly confusing. My preference is to follow published specifications when possible; I
> believe this tends to reduce confusion in the long term. However, if there are strong practical
> reasons to deviate from specifications, that can certainly be done.
> However, we do not want to flip-flop back and forth. To avoid flip-flopping, we need to
> thoroughly understand all the implications of either following the spec or deviating from it. After
> all factors are considered, and as many interested parties as possible are consulted, if the final
> decision is to change, then we should document the reason for the decision and stick to it.
As you said, I think it is correct to follow the published specifications. However, this is a change
that breaks backwards compatibility, and if so, I feel that this change should be discussed in PHP
RFCs and so on.
As about for the convention(customary), at least in Python 3, even if 0x5C or 0x7E is converted to
Shift_JIS, it is converted as it is.
$ python3
Python 3.8.10 (default, Mar 15 2022, 12:22:08)
[GCC 9.4.0] on linux
Type "help", "copyright", "credits" or "license" for
more information.
>>> '\\'.encode("SJIS")
b'\\'
>>> '\\~'.encode("SJIS")
b'\\~'
>>> 'あ'.encode("SJIS")
b'\x82\xa0'
>>>
I tried it with Ruby 2.7. After all it converted as it is.
$ ruby --version
ruby 2.7.0p0 (2019-12-25 revision 647ee6f091) [x86_64-linux-gnu]
$ ruby -e 'puts "あ".encode("SJIS").encode("UTF-8")'
あ
$ ruby -e 'puts "\\".encode("SJIS")'
\
$ ruby -e 'puts "~".encode("SJIS")'
~