Re: BBCodeParser massive update (/me cries)
| From: | bertrand Gugger | Date: | Fri, 14 Oct 2005 12:25:29 +0000 |
| Subject: | Re: BBCodeParser massive update (/me cries) | ||
| References: | 1 2 3 | Groups: | php.pear.dev |
| Request: | Send a blank email to pear-dev+get-40183@lists.php.net to get a copy of this message | ||
Bonjour,
Seth Price wrote:
Hey, I have a chance to look through the Text_Wiki code now. Replying to your comments:I don't guarantee anything, I use the Text_Wiki engine and its Xhtml rendererActually, the whole structure was borked from the beginning, I see no point to reinvent a tree manager to parse BBCode, so it's vain to go ahead in this direction and make the code even more complicated. Text_Wiki 's treeing is implicit.I don't see how you can guarantee XHTML compliant output without some sort of stack based parser.
Simply put, if "[i][b]txt[/i][/b]" results in "<i><b>txt</i></b>" (as it does now), then you can't claim that output will be XHTML. Mismatched BBCode may even qualify as a XSS-style attack because it breaks the validity of your output. You also need some way of guaranteeing that the only tag in <ul> is <li> and other similar PEBKAC errors.I upgraded from CVS including your patch commited by arnaud. I installed it and run the provided example. Simply copied/paste the integrated help. Pehaps your <li> <ul> 's are better Xhtml, but they are _wrong_, adding in all case a first empty element and wrong numbering (the stack :) ? )
The best way to do this is going to be adding a stack based parser in there somewhere. I would propose a "validate" step between the "parse" and "render" steps that can add, remove, and rearrange tokens as needed.Good catch, we need some structure's check before final rendering. There's no best way. I still see no utility in a "stack". We can do that using the existing structure and check the proper nesting.
When you put it there, you can store the results, which is nice because stack based parsing like what needs doing is rather compute intensive.Where there ?
I mean the quoting produced by HTML_BBCodeParser as defined in the ini file: ; possible values: single|double ; use single or double quotes for attributes quotestyle = single From Arnaud's commit:Also you don't respect coding standards and you change some default (quotes).I looked through the PEAR manual and I couldn't find anything against mixing double and single quotes (I may have missed something). I find it produces cleaner code and, last time I checked, using single quotes/concatenation is faster.
- var $_options = array( 'quotestyle' => 'single', + var $_options = array( 'quotestyle' => 'double',
<book title="Advanced PHP Programming" author="George Schlossnagle" page="471"> So although you see a significant improvement in the performance of interpolation [in PHP 4.3], it is still faster to use concatenation to build dynamic strings.</book>Nice :) For sure we miss things (as better Xhtml compliance), but our strength is that any improvment is immediately shared by the other parsers from the package. e.g. if we change Bold in Xhtml renderer to use <strong> instead of <b> (as you point out), all parsers will get it.Also you don't respect coding standards and you change some default (quotes). Your XSS filtering is naturally an enhancement, but it's not complete. Text_Wiki has it from the conception.Yep, all things considered, I like Text_Wiki better, but it is missing some things...
Tests may use what they need to do a better work. Let's see it ...I think Text_Wiki_BBCode would pass the tests the same. (btw it's not in your zip)The tests that I am using are in a wrapper class that would be of little use in PEAR. I'll give you my test cases if you would like them though.
A few things before I go to bed: - The parser seems to add paragraph tags to everything. Processing '"' results in '<p>"</p>', but it should result in '"'. If something should be inline formatted (instead of block), I'd like to keep it that way. (<p> is block formatted by default) How can I get rid of the '<p></p>'?Yes, that's a problem we get, not only in BBCode I fear, I'll look into it. I'm not sure but I think a paragraph should be done only if an empty line is given. Anyway, there's no Raw rule in BBCode so all blocked should be correct.
- You are using a number of non-XHML 1.1 tags such as '<u>', '<i>', and '<b>'. Mind if I replace those with equivalents?I think we can agree on this, perhaps you should get some account/karma ? Note that this is not in the Text_Wiki_BBCode package but in base Text_Wiki Thanks for your interest in this package, I look forward some more cooperation. à+ -- bertrand "toggg" Gugger