[PEPr] Comment on PHP::PHP_LexerGenerator
| From: | Greg Beaver | Date: | Tue, 04 Jul 2006 23:50:35 +0000 |
| Subject: | [PEPr] Comment on PHP::PHP_LexerGenerator | ||
| References: | 1 | Groups: | php.pear.dev |
| Request: | Send a blank email to pear-dev+get-43235@lists.php.net to get a copy of this message | ||
Greg Beaver (http://pear.php.net/user/cellog) has commented on the proposal for
PHP::PHP_LexerGenerator.
Comment:
thanks for trying the package out so extensively Alex :)
1) skip token has been implemented from the beginning, simply use this as
the action:
{return false;}
If you change state and wish to re-process the token, use something like:
{$this->yybegin(self::NEWSTATE);return true;}
To do what is known as "yymore" in flex, use:
{return 'more';}
yymore basically instructs the scanner to ignore the matching rule and try
against the remaining rules.
For your second problem, use a lookahead, like this simple rule:
NODE = /node(?=\s)/
ID = /\w(\w|\d)*/
then, in order to be sure you also match a NODE at the end of input, use:
NODE {}
ID {}
"node" {}
The lexer processes rules in strict order, so that the first rule to match
wins. Many lexers attempt to optimize regular expressions passed in. In
my experience this leads to unpredictable behavior, including the
inability to do fine-grained matching. In short, it makes it impossible
to write a good lexer.
PHP_LexerGenerator allows complete control over the regular expressions
and their speed: you are responsible for optimizing and ordering, but it
will always be possible to look at a lexer .plex file and figure out how
it will work. This is not the case for any other lexer generator (which
is why your example "just works" in other lexer generators).
Hope this explains some of the design choices better.
Proposal information:
http://pear.php.net/pepr/pepr-proposal-show.php?id=415
--
Sent by PEPr, the automatic proposal system at http://pear.php.net