Re: BBCodeParser (transition to Text_Wiki)

From: Date: Fri, 14 Oct 2005 17:12:28 +0000
Subject: Re: BBCodeParser (transition to Text_Wiki)
References: 1 2 3 4  Groups: php.pear.dev 
Request: Send a blank email to pear-dev+get-40193@lists.php.net to get a copy of this message
I don't guarantee anything, I use the Text_Wiki engine and its Xhtml renderer Technically I can't use "<!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.1//EN" "http://www.w3.org/TR/xhtml11/DTD/xhtml11.dtd">" if I don't know if the page is XHTML. If I'm using Text_Wiki in a page, there is no way to tell if the final product is XHTML without running each page through the w3c validator. It would be much easier to fix Text_Wiki to guarantee XHTML output.
I upgraded from CVS including your patch commited by arnaud. I installed it and run the provided example. Simply copied/paste the integrated help. Pehaps your <li> <ul> 's are better Xhtml, but they are _wrong_, adding in all case a first empty element and wrong numbering (the stack :) ? ) I'm not sure if I understand what is going wrong. Can you send me your input and output for this?
Good catch, we need some structure's check before final rendering. There's no best way. I still see no utility in a "stack". We can do that using the existing structure and check the proper nesting. I would be interested in what you have mind for checking proper tag nesting and closure, because I know of no way of doing it without using a stack. (Or you can build a tree, but that is overkill for what we are doing.) Regular expressions simply are theoretically unable to cope with this. We even went over this in my "Intro. to Programming Languages and Compilers" class last year.
Where there ? When you put the validation step in there, you can save the output after validation (but before rendering) to a database for later retrieval. When you need to render the page in XHTML, you can simply retrieve the pre-tokenized and pre-validated text and do the render step. This also allows you to still render out in other formats, but with only the overhead of rendering.
I mean the quoting produced by HTML_BBCodeParser as defined in the ini file: Ah, I didn't even remember making the change.
Tests may use what they need to do a better work. Let's see it ... I attached my tests to the end of my last email, but I don't think that anyone noticed them. I've refined them a bit for BBCodeParser and posted them here:
http://pricepages.org/bbcode/BBCodeParser.phpt.zip ~Seth On Oct 14, 2005, at 7:25 AM, bertrand Gugger wrote:
Bonjour, Seth Price wrote:
Hey, I have a chance to look through the Text_Wiki code now. Replying to your comments:
Actually, the whole structure was borked from the beginning, I see no point to reinvent a tree manager to parse BBCode, so it's vain to go ahead in this direction and make the code even more complicated. Text_Wiki 's treeing is implicit.
I don't see how you can guarantee XHTML compliant output without some sort of stack based parser.
I don't guarantee anything, I use the Text_Wiki engine and its Xhtml renderer
Simply put, if "[i][b]txt[/i][/b]" results in "<i><b>txt</i></b>" (as it does now), then you can't claim that output will be XHTML. Mismatched BBCode may even qualify as a XSS-style attack because it breaks the validity of your output. You also need some way of guaranteeing that the only tag in <ul> is <li> and other similar PEBKAC errors.
I upgraded from CVS including your patch commited by arnaud. I installed it and run the provided example. Simply copied/paste the integrated help. Pehaps your <li> <ul> 's are better Xhtml, but they are _wrong_, adding in all case a first empty element and wrong numbering (the stack :) ? )
The best way to do this is going to be adding a stack based parser in there somewhere. I would propose a "validate" step between the "parse" and "render" steps that can add, remove, and rearrange tokens as needed.
Good catch, we need some structure's check before final rendering. There's no best way. I still see no utility in a "stack". We can do that using the existing structure and check the proper nesting.
When you put it there, you can store the results, which is nice because stack based parsing like what needs doing is rather compute intensive.
Where there ?
Also you don't respect coding standards and you change some default (quotes).
I looked through the PEAR manual and I couldn't find anything against mixing double and single quotes (I may have missed something). I find it produces cleaner code and, last time I checked, using single quotes/concatenation is faster.
I mean the quoting produced by HTML_BBCodeParser as defined in the ini file: ; possible values: single|double ; use single or double quotes for attributes quotestyle = single From Arnaud's commit:
-    var $_options = array(  'quotestyle'    => 'single',
+    var $_options = array(  'quotestyle'    => 'double',
<book title="Advanced PHP Programming" author="George Schlossnagle" page="471"> So although you see a significant improvement in the performance of interpolation [in PHP 4.3], it is still faster to use concatenation to build dynamic strings.</book>
Also you don't respect coding standards and you change some default (quotes). Your XSS filtering is naturally an enhancement, but it's not complete. Text_Wiki has it from the conception.
Yep, all things considered, I like Text_Wiki better, but it is missing some things...
Nice :) For sure we miss things (as better Xhtml compliance), but our strength is that any improvment is immediately shared by the other parsers from the package. e.g. if we change Bold in Xhtml renderer to use <strong> instead of <b> (as you point out), all parsers will get it.
I think Text_Wiki_BBCode would pass the tests the same. (btw it's not in your zip)
The tests that I am using are in a wrapper class that would be of little use in PEAR. I'll give you my test cases if you would like them though.
Tests may use what they need to do a better work. Let's see it ...
A few things before I go to bed: - The parser seems to add paragraph tags to everything. Processing '"' results in '<p>&quot;</p>', but it should result in '&quot;'. If something should be inline formatted (instead of block), I'd like to keep it that way. (<p> is block formatted by default) How can I get rid of the '<p></p>'?
Yes, that's a problem we get, not only in BBCode I fear, I'll look into it. I'm not sure but I think a paragraph should be done only if an empty line is given. Anyway, there's no Raw rule in BBCode so all blocked should be correct.
- You are using a number of non-XHML 1.1 tags such as '<u>', '<i>', and '<b>'. Mind if I replace those with equivalents?
I think we can agree on this, perhaps you should get some account/karma ? Note that this is not in the Text_Wiki_BBCode package but in base Text_Wiki Thanks for your interest in this package, I look forward some more cooperation. à+ --bertrand "toggg" Gugger --PEAR Development Mailing List (http://pear.php.net/) To unsubscribe, visit: http://www.php.net/unsub.php


« previous php.pear.dev (#40193) next »