[PEPr] Comment on PHP::Lexer
| From: | Harry Fuecks | Date: | Tue, 01 Mar 2005 20:46:14 +0000 |
| Subject: | [PEPr] Comment on PHP::Lexer | ||
| References: | 1 | Groups: | php.pear.dev |
| Request: | Send a blank email to pear-dev+get-36443@lists.php.net to get a copy of this message | ||
Harry Fuecks (http://pear.php.net/user/hfuecks) has commented on the proposal for PHP::Lexer.
Comment:
"How were nested tags handled in the Simple Test lexer ?"
Looking at the code again
(http://cvs.sourceforge.net/viewcvs.py/simpletest/simpletest/parser.php?rev=1.66)
the SimpleTest lexer offers an API to end users which is working at a
higher level, the state machine being "bundled" with the lexing
capabilities. You'll notice with methods like
SimpleLexer::addEntryPattern() that it creates new instances of
ParallelRegex - which itself is more or less equivalent to your Lexer. So
I guess someone could build that on top of what you already have.
"is UTF-8 widely used in things susceptible of being parsed by this lexer
?"
By no means an expert on this but here's my take.
Basically we should (as web developers) all be converging on UTF-8 but
parsing, in particular, could become a problem, depending on what your
tokens are, particularily where a specific number of characters is being
searched for.
It's difficult to give a full example here, as the page is encoding as
ISO-8859-1 (Western Europe) but there's some good starting point here:
http://www.intertwingly.net/blog/2005/03/01/Yahoo-Search-and-I18n
For PHP the problem is all the string functions regard a character as
always being a single byte (as it is for the 127 ASCII chars). But in the
example on that blog, some of those characters require multiple bytes to
represent correctly.
That would mean if you had regex that was intended to match sequences of
three characters, seperated by word boundaries like;
/(?<=\b)\w{3}(?=\b)/
You would match 3 ASCII characters but not three multibyte characters. You
could use the /u pattern modifier to instruct the Perl regex engine that
the text is UTF-8 encoded (assuming it is) but in your case you'd probably
also need to be careful using substr() when reducing the remaining text to
parse.
Derick Rethans has some more useful stuff up here:
http://www.derickrethans.nl/files/wereldveroverend-ffm2004.pdf
Also the best general read I've found is
http://www.cs.tut.fi/~jkorpela/chars.html
Proposal information:
http://pear.php.net/pepr/pepr-proposal-show.php?id=197
--
Sent by PEPr, the automatic proposal system at http://pear.php.net