Re: [PEPr] Comment on PHP::Lexer
| From: | Louis Mullie | Date: | Sat, 29 Jan 2005 16:44:54 +0000 |
| Subject: | Re: [PEPr] Comment on PHP::Lexer | ||
| References: | 1 2 | Groups: | php.pear.dev |
| Request: | Send a blank email to pear-dev+get-35786@lists.php.net to get a copy of this message | ||
This is a problem I unfortunately cannot cope with. The problem is that $grammar->addIdent('CHAR', '[a-zA-Z]');
is before $grammar->addIdent('WORD', '{CHAR}+');
and PCRE seems to have problems parsing the longest matching string. The purpose of the _preparePrecedence() method is to cope with that and order the idents based on hierarchy. The new version will not have this problem since it will associate patterns with callbacks (the FLEX way), instead of returning a token stack.
-- Louis
bertrand Gugger wrote:
Hi Louis Mullie:Proposal information: http://pear.php.net/pepr/pepr-proposal-show.php?id=197I come quite right with: <?php require_once 'Lexer.php' ; $grammar = new Lexer_Grammar; $grammar->addIdent('CHAR', '[a-zA-Z]'); $grammar->addIdent('WORD', '{CHAR}+'); $grammar->addIdent('SEP', '\s'); $grammar->addIdent('PARW', '\({WORD}\)'); $lexer = new Lexer($grammar); $stack = $lexer->tokenize('AB A'); print_r($stack); $stack = $lexer->tokenize('A AB'); print_r($stack); $stack = $lexer->tokenize('A (AB)'); print_r($stack); ?> just never recognize AB as Word allways as 2 CHAR :( I would also like to recurse the token within themselves. à+