[PEPr] Comment on PHP::PHP_LexerGenerator
| From: | Alexander Merz | Date: | Mon, 03 Jul 2006 18:44:58 +0000 |
| Subject: | [PEPr] Comment on PHP::PHP_LexerGenerator | ||
| References: | 1 | Groups: | php.pear.dev |
| Request: | Send a blank email to pear-dev+get-43198@lists.php.net to get a copy of this message | ||
Alexander Merz (http://pear.php.net/user/alexmerz) has commented on the proposal for
PHP::PHP_LexerGenerator.
Comment:
Two additional problems:
1.) I implemented a "Skip token" feature in the lexer class:
function advance() {
$ret = $this->{'yylex' .$this->state}();
if($this->token != 17) {
if($ret) {
return true;
}
return false;
} else {
return $this->advance();
}
}
As you can see, it is necessary to know which number the pattern has. It
would be nice to use the pattern name as placeholder. Something like:
if($this->token != %TOKEN%)...
2.) If you define a literal as pattern for a token, you would expect, that
the literal matches "correctly", but this isn't true. So you have trouble
if you have a regex as token pattern, which could also match the literal.
An example:
--- the tokens ---
NODE = "node" (node is a keyword)
ID = /\w(\w|\d)*/ (an Id is just a literal)
------------------
--- The data to parse ---
node [color=grey]
node1 -> node2
-------------------------
In the first line "node" is correctly recognized as NODE. But in the next
line the lexer splits "node1" into "node" and "1". So it is recognized
as
NODE and something other. Thats not correct, it should by ID.
I could solve this by using a regex for NODE instead of a literal:
NODE = /node[^\w\d]/
Other lexer generator does not seem to have such a problem.
Proposal information:
http://pear.php.net/pepr/pepr-proposal-show.php?id=415
--
Sent by PEPr, the automatic proposal system at http://pear.php.net