[php-src] Issue #8281: mb_convert_encoding "\" (backslash) and "~" (tilde) convert failed to Shift_JIS
| From: | zonuexe | Date: | Tue, 05 Apr 2022 16:57:55 +0000 |
| Subject: | [php-src] Issue #8281: mb_convert_encoding "\" (backslash) and "~" (tilde) convert failed to Shift_JIS | ||
| Groups: | php.bugs | ||
| Request: | Send a blank email to php-bugs+get-240673@lists.php.net to get a copy of this message | ||
Issue: https://github.com/php/php-src/issues/8281
Comment Author: zonuexe
Hi @alexdowad. First of all, I would like to express my gratitude and respect to you and the
original developers of mbstring, as I know your refactoring achievements through [this
article](https://qiita.com/rana_kualu/items/61a60ac20fc6583a4af5).
Conversion between multiple character sets is always difficult, and this problem has plagued the
Japanese for 30 years with Unicode, and the conversion map between JIS and Unicode in the original
mbstring is not just a bug in the spec. As the behavior of Ruby shows, it is based on the customs
and use cases of many Japanese users from that time to the present.
unicode.org provides JIS and Unicode conversion maps at <https://www.unicode.org/Public/MAPPINGS/OBSOLETE/EASTASIA/JIS/>,
but this is not part of the standard and is just reference data. Please note in particular.
Replacing
\ with ¥ and preserving it as a backslash are both
"correct" implementations.
The difference between these ideas is also expressed in iconv and
[nkf](https://ja.wikipedia.org/wiki/Network_Kanji_Filter). The Japanese-implemented nkf is not a
standard command, but it is still used by old school Japanese UNIX users.
```
% echo '\~abc' | iconv -f sjis -t utf8
¥‾abc
% echo '\~abc' | nkf -Sw8
\~abc
```
In addition to Wikipedia, several companies and Japanese developers explain how to implement JIS and
Unicode mapping.
* [日本語と文字コード](https://www.kanzaki.com/docs/jcode.html)
* [日本語 Shift-JIS の文字マッピング - IBM
Documentation](https://www.ibm.com/docs/ja/cognos-analytics/10.2.2?topic=SSEP7J_10.2.2/com.ibm.swg.ba.cognos.ug_trnst.10.2.2.doc/c_appendix_round_trip_config.html)
* Same article in English: [Japanese Shift-JIS Character Mapping - IBM
Documentation](https://www.ibm.com/docs/en/cognos-analytics/10.2.2?topic=SSEP7J_10.2.2/com.ibm.swg.ba.cognos.ug_trnst.10.2.2.doc/c_appendix_round_trip_config.html)
* [(プログラマのための)いまさら聞けない標準規格の話 第2回
文字コード実践編 |
オブジェクトの広場](https://www.ogis-ri.co.jp/otc/hiroba/technical/program_standards/part2.html)
* [MySQLのsjisとcp932の違い - tmtms
のメモ](https://blog.tmtms.net/entry/201805/mysql-sjis-cp932)
As mentioned in the discussion so far, Microsoft has burned an extraordinary obsession with
displaying backslash as ¥ in a number of Japanese locales.
Microsoft still describes CP932 to users as Shift_JIS or "シフトJIS", so due to their
efforts many Japanese are unaware of these encodings and character sets. The same is true for many
Japanese PHP programmers.
* [シフトJISで保存 - Microsoft
コミュニティ](https://answers.microsoft.com/ja-jp/msoffice/forum/all/%E3%82%B7%E3%83%95%E3%83%88jis%E3%81%A7%E4%BF%9D/a2caa5f6-52d2-4ca9-992d-a63708c2233f)
This is a post by a forum user, but it contains an Excel screen that says "シフトJIS"
(Shift_JIS).

Although Microsoft has increased Unicode support for Excel in recent years, Japanese users still
believe that converting to Shift_JIS(CP932) for importing and exporting data between the Web and
Excel is a safe and secure method.
* [PHPでExcelで開いても文字化けしないCSVを出力する -
Qiita](https://qiita.com/ikemonn/items/f2bc4f9f834c989084ff)
* [Excelで文字が自動変換されないcsvをPHP(Laravel)で出力する技術 -
テコテック開発者ブログ](https://tec.tecotec.co.jp/entry/2020/12/12/000000)
* [PHP:
fgetcsvでもSJISのCSVをUTF-8として《安全》に読む方法(ストリームフィルタ使用)
- Qiita](https://qiita.com/suin/items/3edfb9cb15e26bffba11)
These are just a few, and some of these articles show code that outputs broken CSV, but many
Japanese users are more concerned about "文字化け".
It is believed that many Japanese companies using Windows still use that method. Unfortunately, many
of them are not interested in disseminating information to the tech community. (Moreover, many
programmers hired by such companies may not be aware of Composer's existence...)
----
I think changing the character encoding and conversion map should have been a careful debate, but I
agree that "flip-flops" can cause further confusion.
Converting all SJIS to CP932 is not a good option for backwards
compatibility as Shift_JIS and CP932 have different conversion maps for Chinese characters.
The improvement I suggest is to specify in [PHP: mb_convert_encoding -
Manual](https://www.php.net/mb_convert_encoding) that the conversion map has changed in PHP 8.1 and
provide a backwards compatible and secure workaround.
Since the only characters converted from ASCII in PHP 8.1 are ~ and \
(<https://3v4l.org/eLeHE>), it is possible to maintain
compatibility of the conversion results by converting these with strtr().
```php
$str = strtr(mb_convert_encoding($str, 'UTF-8', 'SJIS'), ['¥' =>
'\\', '‾' => '~']);
```
I am grateful to all of you for your efforts on these issues.