Re: Text_Wiki: a new rendering algorithm

From: Date: Sun, 05 Mar 2006 14:06:59 +0000
Subject: Re: Text_Wiki: a new rendering algorithm
References: 1  Groups: php.pear.dev 
Request: Send a blank email to pear-dev+get-41662@lists.php.net to get a copy of this message
Justin Patrin wrote:
toggg pointed out a webpage (http://d.hatena.ne.jp/sumipan/20060304/1141493832) which seems to point out that Text_Wiki is doing char-by-char iteration in order to render the tokens in the parsed source. I'd of course seen this before, but hadn't really thought about it. I implemented a simple preg_replace_callback solution as well and it seems to be around 20% faster. I'm not sure about memory usage, however. I've left in the old code as well. To use the new algorithm set the newRendering member variable on the object. If people could test the new algorithm and possibly do some memory and CPU tests of their own I'd appreciate it. (Or if anyone has any suggestions for how to profile the code with free tools I'll try to test it more myself.) -- Justin Patrin Cool try you did of it, Justin.
So far I can see, that should not change anything, just lower resources comsume :) (anyway, you made it experimental thru this necessary $wiki->newRendering, nice) Not sure what that does if renders change $this->wiki->source, but problem is same (or worse) in the classical solution, source should be kinda protected there, e.g. we should work on a copy (hehe, perhaps implicit in preg_replace_callback() :) ) I just checked the sources, did not yet run tests... Actually, as this thing is growing, we need "unit" tests. Example, we should be able to take your change and run it thru... As any very versatile / configurable package (don't say MDB2) , we have a lot of different configurations / items to check. I see the stuff as * (source , parser) |
                   |-P-> tokenized |
    * parserConfig |               |
                        * renderer |-R-> result
                                   |
                  * rendererConfig |
where: "*" means "implementer/user/ input and *source* is the text to transform, I mean (source , parser) considering the syntax is inherent to it. (current doc/Test_Text_Wiki.php tryes to guess it from source) *parser* is Default, Cowiki, Dokuwiki, Tiki, BBCode, Mediawiki ... *parserConfig* is its configuration, may be empty then default config for this parser *tokenized* is the common tokenized form, transformed source text with the indexed tokens pointing to the associated parsed option array *renderer* is the choosen renderer, often default to Xhtml *rendererConfig* is its configuration (global and for each rules), may be empty then default for the renderer *result* what we expect We have to test/assert * -P-> : the parsers, internal tokenized form: "remaining" source with tokens and them options array could be asserted thru var_dump() * -R-> : the renderers using tokenized form, could be asserted thru EXPECTF or EXPECTREGEX from pear's .phpt I'm normally opposed to tests based on some "internal" states, anyway, I'm considering we should taylor the tests and not mix problems, so 2 separated sets of tests (belonging to each concerned sub-package) would be nice. I intend to have it the most simple and natural possible, then pear's .phpt are only possible. I would like to cross / re-use source and parserConfig in combinations tokenized, renderer and rendereConfig the same. Then I guess we need some kind of test generator. Again talking, again bad californian slang :) Regards -- toggg

« previous php.pear.dev (#41662) next »