Re: BBCodeParser massive update (/me cries)

From: Date: Fri, 14 Oct 2005 16:58:04 +0000
Subject: Re: BBCodeParser massive update (/me cries)
References: 1 2 3 4  Groups: php.pear.dev 
Request: Send a blank email to pear-dev+get-40192@lists.php.net to get a copy of this message
On 10/14/05, bertrand Gugger <bertrand@toggg.com> wrote: > Bonjour, > Seth Price wrote: > Responding more to Seth than bertrand... > > Hey, I have a chance to look through the Text_Wiki code now. Replying > > to your comments: > > > >> Actually, the whole structure was borked from the beginning, I see > >> no point to reinvent a tree manager to parse BBCode, so it's vain to > >> go ahead in this direction and make the code even more complicated. > >> Text_Wiki 's treeing is implicit. > > > > I don't see how you can guarantee XHTML compliant output without some > > sort of stack based parser. > > I don't guarantee anything, I use the Text_Wiki engine and its Xhtml > renderer Garbage In, Garbage Out! The XHtml renderer outputs XHtml if you give it good input. If you're worried about it I suggest you run the output through tidy (or similar) upon submission and reject it if the XML is broken. It doesn't have to be validated every time you render it, only when it's submitted. This leads me to believe it's outside the scope of Text_Wiki. > > > Simply put, if "[i][b]txt[/i][/b]" results in > > "<i><b>txt</i></b>" (as > > it does now), then you can't claim that output will be XHTML. > > Mismatched BBCode may even qualify as a XSS-style attack because it > > breaks the validity of your output. Same as above. I don't agree that this is an XSS attack. An XSS attack injects unwanted code into yout page. One thing which *could* be construed as an XSS attack would be if a parser allows unmatched tags like [i] to be rendered without an ending [/i]. I don't *think* that any of our parsers allow this but if they do, submit a bug report. > > > > You also need some way of guaranteeing that the only tag in <ul> is > > <li> and other similar PEBKAC errors. Text_Wiki *shouldn't* be putting anything inside lists. Then again, I'm used to wiki markup which doesn't normally have a "start list" and an "end list" tag. Yep, upon looking it seems that BBCode has a [list] [/list] format. This is simply bad syntax on the part of BBCode. BBCode is the thing allowing other "tags" within the [list] syntax, not Text_Wiki. If you compare, for example, TikiWiki syntax: *One *Two **Two-sub There is no way to put other things in a list. It begins and ends when the list items stop. > > I upgraded from CVS including your patch commited by arnaud. I installed > it and run the provided example. Simply copied/paste the integrated > help. Pehaps your <li> <ul> 's are better Xhtml, but they are _wrong_, > adding in all case a first empty element and wrong numbering (the stack > :) ? ) > > > The best way to do this is going to be adding a stack based parser in > > there somewhere. I would propose a "validate" step between the > > "parse" and "render" steps that can add, remove, and rearrange tokens > > as needed. > > Good catch, we need some structure's check before final rendering. > There's no best way. I still see no utility in a "stack". We can do that > using the existing structure and check the proper nesting. > > > > > When you put it there, you can store the results, which is nice > > because stack based parsing like what needs doing is rather compute > > intensive. > Text_Wiki is only a text processor. It converts from one format to another. If you give it bad input you're going to get bad output. Text_Wiki itself knows nothing about validity of structures. It *could* be possible to validate the input....possibly....but this would be something that should be done as little as possible due to the extra overhead. If you want to see an interesting solution for fixing the XML-ness of HTML, here's a snippet of code I wrote: http://pear.reversefold.com/fixHtml.php It makes sure that all XML tags are ended/started correctly and even adds more tags to (hopefully) keep the original intent (<b>bold<i>italic bold</b>italic<i> becomes <b>bold<i>italic bold</i></b><i>italic<i> for example). It's not perfect, but it does fix XML validity at least. Now I understand the "need" for verifying syntax and output, but I can see only one way that Text_Wiki can conceivably do this. First of all, I'll restate that I think this should be *optional* and off by default. Verification should only be done upon input and not on normal output, otherwise you're going to be verifying the same thing over and over again. (This is, of course, an application-level decision, which is why I'm stressing the need for flexibility.) Second, this verification should probably be done on the intermediate format. This way the verification need not be specific to any input or output syntax. For this to work, we (the Text_Wiki devs) will have to alter some of the "tokens" to be more standard. To verify the "XML-ness" of the tokens we need to have the 'type' key for start and end tokens be the same and have a key which is always 'start' or 'end' so that we can easily match them *without* knowing specifics about the tokens. This was we can use a simple stack-based approach for validation of start/close tags. This should take care of validating start/end tokens. The harder part is validation of XHtml structures. I don't think that this belongs in Text_Wiki at all. It's simply out of scope. Text_Wiki doesn't really know anything about the input and output formats, it only converts them (have I said this before?). We *could* add an optional validator, say, for the different input syntaxes, but I have no idea how this could be sanely implemented. IMHO this should be left to the application to do. If you're that worried about XHtml compliance (I try not to when I'm letting users edit things. They will *never* all understand what XHtml is and why they should enter things a certain way.) then the output should be run through tidy or some other validation engine and rejected if it is incorrect. Or you could allow tidy to fix it (make it XHtml) and cache *that* as the output. -- Justin Patrin

« previous php.pear.dev (#40192) next »