Re: BBCodeParser (transition to Text_Wiki)

From: Date: Fri, 14 Oct 2005 18:12:51 +0000
Subject: Re: BBCodeParser (transition to Text_Wiki)
References: 1 2 3 4 5  Groups: php.pear.dev 
Request: Send a blank email to pear-dev+get-40195@lists.php.net to get a copy of this message
Garbage In, Garbage Out! The XHtml renderer outputs XHtml if you give it good input. If you're worried about it I suggest you run the output through tidy (or similar) upon submission and reject it if the XML is broken. It doesn't have to be validated every time you render it, only when it's submitted. This leads me to believe it's outside the scope of Text_Wiki. But it doesn't have to be like this. I'm aiming at supporting code written by people that don't know XML. I don't want to force people to learn proper XHTML before they submit formatted text with my site. If I can do it with BBCodeParser, then we can do it with Text_Wiki. With my current design, it is also only validated when submitted.
Same as above. I don't agree that this is an XSS attack. An XSS attack injects unwanted code into yout page. One thing which *could* be construed as an XSS attack would be if a parser allows unmatched tags like [i] to be rendered without an ending [/i]. I don't *think* that any of our parsers allow this but if they do, submit a bug report. One of the purposes for XHTML is that it can be rendered by an lightweight XML engine. If it doesn't match the DTD, then we have an error. Most web browsers are smarter than just displaying a rendering error message, but if someone can input text that forces the browser to correct the XHTML or error out, then that person has broken your page. Technically, it would be perfectly valid for the browser do display only a XML rendering error message.
Yep, upon looking it seems that BBCode has a [list] [/list] format. This is simply bad syntax on the part of BBCode. BBCode is the thing allowing other "tags" within the [list] syntax, not Text_Wiki. If you compare, for example, TikiWiki syntax: XHTML requires <ul> and </ul> tags, but I don't consider that "bad syntax on the part of" XHTML.
If you want to see an interesting solution for fixing the XML-ness of HTML, here's a snippet of code I wrote: http://pear.reversefold.com/fixHtml.php That's starting to look similar to the stack based engine in HTML_BBCodeParser ;)
Now I understand the "need" for verifying syntax and output, but I can [snip] Yep, this is the same thought process that I have had, but I still want to work in a validation engine for XHTML structures somewhere. Part of me thinks that this should be built into the BBCode parser, because:
[*] list item should be turned into "proper" BBCode: [list][li]list item[/li][/list] But then I'm sure that the various Wiki codes could benefit from proper XHTML structure validation also, so maybe it should be worked into the rendering engine. But then every page retrieval needs to re-correct the broken BBCode. And I'm not sure how Latex works, but I'm sure that it could use some structure validation. So it could go in either place. Maybe we should implement a parser that can read XHTML, so then I can render it as XHTML with the structure validation, and then parse it back into intermediate format again for storage. It would look like this: -(messy BBCode) [*] list item 1. parse the BBCode and do tag matching in a stack -(intermediate format with improper structure) {li token}list item{/li token} 2. render as XHTML and do structure validation (most likely in another stack) -(XHTML 1.1 compliant) <ul><li>list item</li></ul> 3. parse the XHTML -(intermediate format with proper structure for XHTML) {list token}{li token}list item{/li token}{/list token} 4. store in database This would be a royal pain the the butt to implement, and would require both a stack based parser and a stack based render, but afterwards would have the best and most extensible Wiki/BBCode/XHTML parser/render available for use. Again, I don't know Latex, but the same method could be used to validate anything intended for Latex output with the addition of a Latex parser. You could render XHTML, parse XHTML, render Latex, parse Latex, and have intermediate format that is *guaranteed* valid in *both* Latex and XHTML. Oh... one can dream... Personally, I can't find a package out there that does a quality job rendering BBCode. I think the closest now is HTML_BBCodeParser after the patches that I made, but it still really isn't *that* good. I was just telling my aunt that getting a computer to understand a human is just as hard as getting a human to understand a computer. She seemed to understand. ;) ~Seth On Oct 14, 2005, at 11:58 AM, Justin Patrin wrote:
On 10/14/05, bertrand Gugger <bertrand@toggg.com> wrote:
Bonjour, Seth Price wrote:
Responding more to Seth than bertrand...
Hey, I have a chance to look through the Text_Wiki code now. Replying to your comments:
Actually, the whole structure was borked from the beginning, I see no point to reinvent a tree manager to parse BBCode, so it's vain to go ahead in this direction and make the code even more complicated. Text_Wiki 's treeing is implicit.
I don't see how you can guarantee XHTML compliant output without some sort of stack based parser.
I don't guarantee anything, I use the Text_Wiki engine and its Xhtml renderer
Garbage In, Garbage Out! The XHtml renderer outputs XHtml if you give it good input. If you're worried about it I suggest you run the output through tidy (or similar) upon submission and reject it if the XML is broken. It doesn't have to be validated every time you render it, only when it's submitted. This leads me to believe it's outside the scope of Text_Wiki.
Simply put, if "[i][b]txt[/i][/b]" results in "<i><b>txt</i></b>" (as it does now), then you can't claim that output will be XHTML. Mismatched BBCode may even qualify as a XSS-style attack because it breaks the validity of your output.
Same as above. I don't agree that this is an XSS attack. An XSS attack injects unwanted code into yout page. One thing which *could* be construed as an XSS attack would be if a parser allows unmatched tags like [i] to be rendered without an ending [/i]. I don't *think* that any of our parsers allow this but if they do, submit a bug report.
You also need some way of guaranteeing that the only tag in <ul> is <li> and other similar PEBKAC errors.
Text_Wiki *shouldn't* be putting anything inside lists. Then again, I'm used to wiki markup which doesn't normally have a "start list" and an "end list" tag. Yep, upon looking it seems that BBCode has a [list] [/list] format. This is simply bad syntax on the part of BBCode. BBCode is the thing allowing other "tags" within the [list] syntax, not Text_Wiki. If you compare, for example, TikiWiki syntax: *One *Two **Two-sub There is no way to put other things in a list. It begins and ends when the list items stop.
I upgraded from CVS including your patch commited by arnaud. I installed it and run the provided example. Simply copied/paste the integrated help. Pehaps your <li> <ul> 's are better Xhtml, but they are _wrong_, adding in all case a first empty element and wrong numbering (the stack :) ? )
The best way to do this is going to be adding a stack based parser in there somewhere. I would propose a "validate" step between the "parse" and "render" steps that can add, remove, and rearrange tokens as needed.
Good catch, we need some structure's check before final rendering. There's no best way. I still see no utility in a "stack". We can do that using the existing structure and check the proper nesting.
When you put it there, you can store the results, which is nice because stack based parsing like what needs doing is rather compute intensive.
Text_Wiki is only a text processor. It converts from one format to another. If you give it bad input you're going to get bad output. Text_Wiki itself knows nothing about validity of structures. It *could* be possible to validate the input....possibly....but this would be something that should be done as little as possible due to the extra overhead. If you want to see an interesting solution for fixing the XML-ness of HTML, here's a snippet of code I wrote: http://pear.reversefold.com/fixHtml.php It makes sure that all XML tags are ended/started correctly and even adds more tags to (hopefully) keep the original intent (<b>bold<i>italic bold</b>italic<i> becomes <b>bold<i>italic bold</i></b><i>italic<i> for example). It's not perfect, but it does fix XML validity at least. Now I understand the "need" for verifying syntax and output, but I can see only one way that Text_Wiki can conceivably do this. First of all, I'll restate that I think this should be *optional* and off by default. Verification should only be done upon input and not on normal output, otherwise you're going to be verifying the same thing over and over again. (This is, of course, an application-level decision, which is why I'm stressing the need for flexibility.) Second, this verification should probably be done on the intermediate format. This way the verification need not be specific to any input or output syntax. For this to work, we (the Text_Wiki devs) will have to alter some of the "tokens" to be more standard. To verify the "XML-ness" of the tokens we need to have the 'type' key for start and end tokens be the same and have a key which is always 'start' or 'end' so that we can easily match them *without* knowing specifics about the tokens. This was we can use a simple stack-based approach for validation of start/close tags. This should take care of validating start/end tokens. The harder part is validation of XHtml structures. I don't think that this belongs in Text_Wiki at all. It's simply out of scope. Text_Wiki doesn't really know anything about the input and output formats, it only converts them (have I said this before?). We *could* add an optional validator, say, for the different input syntaxes, but I have no idea how this could be sanely implemented. IMHO this should be left to the application to do. If you're that worried about XHtml compliance (I try not to when I'm letting users edit things. They will *never* all understand what XHtml is and why they should enter things a certain way.) then the output should be run through tidy or some other validation engine and rejected if it is incorrect. Or you could allow tidy to fix it (make it XHtml) and cache *that* as the output. -- Justin Patrin -- PEAR Development Mailing List (http://pear.php.net/) To unsubscribe, visit: http://www.php.net/unsub.php


« previous php.pear.dev (#40195) next »