RE: [PHP] detecting 2-bit characters.

From: Date: Wed, 15 Nov 2000 03:51:31 +0000
Subject: RE: [PHP] detecting 2-bit characters.
Groups: php.general 
Request: Send a blank email to php-general+get-25341@lists.php.net to get a copy of this message
I probably have the same issue (or soon will do) ;) My thinking is to do a search for higher ascii values eg ascii chars >128 Chinese / Japanese etc are double byte (is hiragana?)/katakana is (or was it the other way around...) most of the double byte pairs include higher order ascii value. Problems with this approach - French and some european characters also have higher ascii values - the ecrit, the umlaut etc. This may or may not be an issue for you. Detecting the language capabilities of the browser doesn't really help - in China for example, you can run english os, english exploder, plus chinese star or similar ime, and the browser returns US lang type. You may want to have a look at the unicode functions in php as well utf8_decode / encode might help - i haven't played with them yet to see if utf8_decode on a non ascii string will modify it or not. so comparison checks with unicode vs iso US-ASCII decoded -> encoded are different may work based on the text below from the manual (if you're unsure what i'm getting at what I mean is) unicode ÄãºÃ should turn into ?? in ascii conversion, so converting a string returned from the browser down to the lowest denominator, then comparing with the original string should show if its ascii or not. If its ascii then it won't change, if its unicode double byte, then it will change the non ascii to question marks From the manual: Target encoding is done when PHP passes data to XML handler functions. When an XML parser is created, the target encoding is set to the same as the source encoding, but this may be changed at any point. The target encoding will affect character data as well as tag names and processing instruction targets. If the XML parser encounters characters outside the range that its source encoding is capable of representing, it will return an error. If PHP encounters characters in the parsed XML document that can not be represented in the chosen target encoding, the problem characters will be "demoted". Currently, this means that such characters are replaced by a question mark. -----Original Message----- From: Maxim Maletsky [mailto:maxim.maletsky@japaninc.net] Sent: November 15, 2000 11:32 AM To: 'PHP General List. (E-mail)' Subject: [PHP] detecting 2-bit characters. Hello, Our submission form is being submitted 30% by Japanese nationals and they (stupidly even if we mention it) write their data in Japanese instead of English. (the product and the whole website is in English). It is now to time to disallow it. q: How can I detect if a string is contains over 1-bit characters ? I do have an idea how to write this kind of thing but I was wondering if there's a build-in function in PHP. Anyone used/wrote anything like this? Just to detect this for me is enough. Thanks, it would save lots of my time. Maxim Maletsky - maxim@j-door.com <mailto:maxim@j-door.com> Webmaster, J-Door.com / J@pan Inc. LINC Media, Inc. TEL: 03-3499-2175 x 1271 FAX: 03-3499-3109 http://www.j-door.com <http://www.j-door.com/> http://www.japaninc.net <http://www.japaninc.net/> http://www.lincmedia.co.jp <http://www.lincmedia.co.jp/>

« previous php.general (#25341) next »