[PEPr] Comment on PHP::Lexer

From: Date: Tue, 01 Mar 2005 20:46:14 +0000
Subject: [PEPr] Comment on PHP::Lexer
References: 1  Groups: php.pear.dev 
Request: Send a blank email to pear-dev+get-36443@lists.php.net to get a copy of this message
Harry Fuecks (http://pear.php.net/user/hfuecks) has commented on the proposal for PHP::Lexer. Comment: "How were nested tags handled in the Simple Test lexer ?" Looking at the code again (http://cvs.sourceforge.net/viewcvs.py/simpletest/simpletest/parser.php?rev=1.66) the SimpleTest lexer offers an API to end users which is working at a higher level, the state machine being "bundled" with the lexing capabilities. You'll notice with methods like SimpleLexer::addEntryPattern() that it creates new instances of ParallelRegex - which itself is more or less equivalent to your Lexer. So I guess someone could build that on top of what you already have. "is UTF-8 widely used in things susceptible of being parsed by this lexer ?" By no means an expert on this but here's my take. Basically we should (as web developers) all be converging on UTF-8 but parsing, in particular, could become a problem, depending on what your tokens are, particularily where a specific number of characters is being searched for. It's difficult to give a full example here, as the page is encoding as ISO-8859-1 (Western Europe) but there's some good starting point here: http://www.intertwingly.net/blog/2005/03/01/Yahoo-Search-and-I18n For PHP the problem is all the string functions regard a character as always being a single byte (as it is for the 127 ASCII chars). But in the example on that blog, some of those characters require multiple bytes to represent correctly. That would mean if you had regex that was intended to match sequences of three characters, seperated by word boundaries like; /(?<=\b)\w{3}(?=\b)/ You would match 3 ASCII characters but not three multibyte characters. You could use the /u pattern modifier to instruct the Perl regex engine that the text is UTF-8 encoded (assuming it is) but in your case you'd probably also need to be careful using substr() when reducing the remaining text to parse. Derick Rethans has some more useful stuff up here: http://www.derickrethans.nl/files/wereldveroverend-ffm2004.pdf Also the best general read I've found is http://www.cs.tut.fi/~jkorpela/chars.html Proposal information: http://pear.php.net/pepr/pepr-proposal-show.php?id=197 -- Sent by PEPr, the automatic proposal system at http://pear.php.net

« previous php.pear.dev (#36443) next »