RE: [PHP] detecting 2-bit characters.
| From: | Lawrence dot Sheed at dfait-maeci dot gc dot ca | Date: | Wed, 15 Nov 2000 04:17:39 +0000 |
| Subject: | RE: [PHP] detecting 2-bit characters. | ||
| Groups: | php.general | ||
| Request: | Send a blank email to php-general+get-25348@lists.php.net to get a copy of this message | ||
There are 256 characters in ascii correct.
7 bits of those are used for lower ascii values - eg a-z, A-Z, !@#$%^%*()-=
(etc)
the upper ascii characters 128 - 256 are different for each code set you are
using.
English text without any funny stuff like è or é or à or ê uses characters
<128.
A quick and dirty check would be to see if the following 4 bytes/ 2 double
byte characters such as ?? (GB Encoded RI BEN or japan in chinese) are
upper ascii (which they generally are).
PHP isn't going to know what encoding they've used natively. What i'm
suggesting is to either implement a quick and dirty fix like xor'ing the
character with 128 (or shifting up 128, and back down again) and seeing if
it changes.
The other way is to ut8_decode to us-ascii, and see if the string stays the
same.
according to the manual, non ascii characters will be decoded to ? marks.
so a check something like the following:
$newstring = utf8_decode ($string);
if ($newstring == $string) {
//isnt iso us-ascii
}
else {
//is iso us-ascii
}
would probably do it. Note that the above won't work probably because we
haven't specified character sets to decode to.
Hope this is clearer.
Cheers,
Lawrence
-----Original Message-----
From: Maxim Maletsky [mailto:maxim.maletsky@japaninc.net]
Sent: November 15, 2000 12:04 PM
To: 'Lawrence.Sheed@dfait-maeci.gc.ca'
Cc: php-general@lists.php.net
Subject: RE: [PHP] detecting 2-bit characters.
there are 256 ASCII characters each of them is one bit only.
so only these are OK. it will contain then all the European languages even
French and Italian. But not Russian, Chinese and Japanese (nor all these
Arabian and Indian and Korean ect,, as well..)
so let's say I chunk_split() the word, how then can I say that the character
A is one bit and ? is a 2-bit one?
That is making me assume that it could be a built-in function in PHP that
knows how many bits are in a character.
Thanks,
Maxim Maletsky
-----Original Message-----
From: Lawrence.Sheed@dfait-maeci.gc.ca
[mailto:Lawrence.Sheed@dfait-maeci.gc.ca]
Sent: Wednesday, November 15, 2000 12:52 PM
To: Maxim Maletsky
Cc: php-general@lists.php.net
Subject: RE: [PHP] detecting 2-bit characters.
I probably have the same issue (or soon will do) ;)NI
My thinking is to do a search for higher ascii values eg ascii chars >128
Chinese / Japanese etc are double byte (is hiragana?)/katakana is (or was it
the other way around...)
most of the double byte pairs include higher order ascii value.
Problems with this approach - French and some european characters also have
higher ascii values - the ecrit, the umlaut etc. This may or may not be an
issue for you.
Detecting the language capabilities of the browser doesn't really help - in
China for example, you can run english os, english exploder, plus chinese
star or similar ime, and the browser returns US lang type.
You may want to have a look at the unicode functions in php as well
utf8_decode / encode might help - i haven't played with them yet to see if
utf8_decode on a non ascii string will modify it or not. so comparison
checks with unicode vs iso US-ASCII decoded -> encoded are different may
work based on the text below from the manual (if you're unsure what i'm
getting at what I mean is)
unicode ÄãºÃ should turn into ?? in ascii conversion, so converting a
string returned from the browser down to the lowest denominator, then
comparing with the original string should show if its ascii or not. If its
ascii then it won't change, if its unicode double byte, then it will change
the non ascii to question marks
From the manual:
Target encoding is done when PHP passes data to XML handler functions. When
an XML parser is created, the target encoding is set to the same as the
source encoding, but this may be changed at any point. The target encoding
will affect character data as well as tag names and processing instruction
targets.
If the XML parser encounters characters outside the range that its source
encoding is capable of representing, it will return an error.
If PHP encounters characters in the parsed XML document that can not be
represented in the chosen target encoding, the problem characters will be
"demoted". Currently, this means that such characters are replaced by a
question mark.
-----Original Message-----
From: Maxim Maletsky [mailto:maxim.maletsky@japaninc.net]
Sent: November 15, 2000 11:32 AM
To: 'PHP General List. (E-mail)'
Subject: [PHP] detecting 2-bit characters.
Hello,
Our submission form is being submitted 30% by Japanese nationals and they
(stupidly even if we mention it) write their data in Japanese instead of
English. (the product and the whole website is in English). It is now to
time to disallow it.
q: How can I detect if a string is contains over 1-bit characters ?
I do have an idea how to write this kind of thing but I was wondering if
there's a build-in function in PHP.
Anyone used/wrote anything like this?
Just to detect this for me is enough.
Thanks, it would save lots of my time.
Maxim Maletsky - maxim@j-door.com <mailto:maxim@j-door.com>
Webmaster, J-Door.com / J@pan Inc.
LINC Media, Inc.
TEL: 03-3499-2175 x 1271
FAX: 03-3499-3109
http://www.j-door.com <http://www.j-door.com/>
http://www.japaninc.net <http://www.japaninc.net/>
http://www.lincmedia.co.jp <http://www.lincmedia.co.jp/>
--
PHP General Mailing List (http://www.php.net/)
To unsubscribe, e-mail: php-general-unsubscribe@lists.php.net
For additional commands, e-mail: php-general-help@lists.php.net
To contact the list administrators, e-mail: php-list-admin@lists.php.net