tokenize words with foreign language chars problem

From: Date: Tue, 01 Aug 2000 20:38:59 +0000
Subject: tokenize words with foreign language chars problem
Groups: php.general 
Request: Send a blank email to php-general+get-9547@lists.php.net to get a copy of this message
Hi All ! My application receives words (may be separated with + or - chars) to do a search into the database. I need to split the string I received to construct the database query string, but I've many errors when people insert words with accents, etc (for example in Spanish language). To tokenize I'm using the function $token = preg_split("/[^\w\+\-]/", $search_word); with the problem related above. Also I've tried to use rawurlencode function previous to call preg_split, but if people insert the word "Clarín" after rwaurlencode I've the word "Clar%EDn" and after preg_split I 've 2 words "Clar" and "EDn" !!! How can I solve the problem ? Any help will be appreciated... Javier

« previous php.general (#9547) next »