Re: [PEPr] Comment on PHP::Lexer
| From: | bertrand Gugger | Date: | Sat, 29 Jan 2005 01:47:22 +0000 |
| Subject: | Re: [PEPr] Comment on PHP::Lexer | ||
| References: | 1 2 3 | Groups: | php.pear.dev |
| Request: | Send a blank email to pear-dev+get-35777@lists.php.net to get a copy of this message | ||
OK, Louis, time is no key ;)
Did you consider the other points in my comment: cyclic, FSM pckg;
and later in dev-list '(?' in regexp ?
Perhaps you should do regexp on regexp ?
A proposal is rough, give rough examples, not tokenize php.
For the line stated, it's an obvious mistake, the ", $rule" at the end should be removed (corrected version @ http://www.mulliemedia.com/lexer/Lexer2.php). Anyways I am completely refactoring this, I should have a 2x better version running by next weekend, but until then here is a working example (sorry for the longness) :
<?php
require_once 'Lexer.php';
$grammar = new Lexer_Grammar();
try {
/**
* Comments have the highest precedence
**/
$grammar->addIdent('COMMENT', '\/\/.*', LEXER_IGNORE);
/**
* A literal string. For simplicity, escaping " in a literal
* is not permitted.
**/
$grammar->addIdent('LITERAL', '".*?"');
/**
* End of statement
**/
$grammar->addIdent('END_STATEMENT', ';');
/**
* Operators
**/
$grammar->addIdent('IS_EQUAL', '==');
$grammar->addIdent('ASSIGN', '=');
$grammar->addIdent('CONCAT', '\.');
/**
* Structures
**/
$grammar->addIdent('LEFT_PARENTHESIS', '\(');
$grammar->addIdent('RIGHT_PARENTHESIS', '\)');
$grammar->addIdent('LEFT_BRACKET', '\{');
$grammar->addIdent('RIGHT_BRACKET', '\}');
/**
* Functions
**/
$grammar->addIdent('PRINT', '[Pp][Rr][Ii][Nn][Tt]');
/**
* Control structures
**/
$grammar->addIdent('IF', '[Ii][Ff]');
$grammar->addIdent('VAR_DECL', '[Vv][Aa][Rr]');
/**
* Spaces
**/
$grammar->addIdent('SPACE', '\\s+', LEXER_IGNORE);
/**
* Other input
**/
$grammar->addIdent('CHAR', '[a-zA-Z]');
$grammar->addIdent('INT', '[0-9]+');
$grammar->addIdent('VARIABLE', '\\${CHAR}+');
/**
* Throw an error for any character that wouldn't be matched
**/
$grammar->setOption(LEXER_IGNORE_NON_MATCHED, false);
}
catch (Lexer_Exception $e) {
die($e->getMessage());
}
$lexer = new Lexer($grammar);
$start = microtime(true);
try {
$program = <<<EOP
// Variable declaration
var \$var;
\$var = 1;
if (\$var == 1) {
print("Hello, world ! The \$var variable //has a value of int(1). The literal text" .
"should not be parsed.");
}
EOP;
$stack = $lexer->tokenize($program);
} catch (Lexer_Exception $e) {
die($e->getMessage());
}
$end = microtime(true);
$time = $end - $start;
print '<br><br><br>';
print 'Exec time : ' . $time .' Tokens scanned : ' . count($stack) . '<br /> Stack :';
print_r($stack);
?>