Re: BBCodeParser massive update (/me cries)
| From: | Justin Patrin | Date: | Fri, 14 Oct 2005 16:58:04 +0000 |
| Subject: | Re: BBCodeParser massive update (/me cries) | ||
| References: | 1 2 3 4 | Groups: | php.pear.dev |
| Request: | Send a blank email to pear-dev+get-40192@lists.php.net to get a copy of this message | ||
On 10/14/05, bertrand Gugger <bertrand@toggg.com> wrote:
> Bonjour,
> Seth Price wrote:
>
Responding more to Seth than bertrand...
> > Hey, I have a chance to look through the Text_Wiki code now. Replying
> > to your comments:
> >
> >> Actually, the whole structure was borked from the beginning, I see
> >> no point to reinvent a tree manager to parse BBCode, so it's vain to
> >> go ahead in this direction and make the code even more complicated.
> >> Text_Wiki 's treeing is implicit.
> >
> > I don't see how you can guarantee XHTML compliant output without some
> > sort of stack based parser.
>
> I don't guarantee anything, I use the Text_Wiki engine and its Xhtml
> renderer
Garbage In, Garbage Out! The XHtml renderer outputs XHtml if you give
it good input. If you're worried about it I suggest you run the output
through tidy (or similar) upon submission and reject it if the XML is
broken. It doesn't have to be validated every time you render it, only
when it's submitted. This leads me to believe it's outside the scope
of Text_Wiki.
>
> > Simply put, if "[i][b]txt[/i][/b]" results in
> > "<i><b>txt</i></b>" (as
> > it does now), then you can't claim that output will be XHTML.
> > Mismatched BBCode may even qualify as a XSS-style attack because it
> > breaks the validity of your output.
Same as above. I don't agree that this is an XSS attack. An XSS attack
injects unwanted code into yout page. One thing which *could* be
construed as an XSS attack would be if a parser allows unmatched tags
like [i] to be rendered without an ending [/i]. I don't *think* that
any of our parsers allow this but if they do, submit a bug report.
> >
> > You also need some way of guaranteeing that the only tag in <ul> is
> > <li> and other similar PEBKAC errors.
Text_Wiki *shouldn't* be putting anything inside lists. Then again,
I'm used to wiki markup which doesn't normally have a "start list" and
an "end list" tag.
Yep, upon looking it seems that BBCode has a [list] [/list] format.
This is simply bad syntax on the part of BBCode. BBCode is the thing
allowing other "tags" within the [list] syntax, not Text_Wiki. If you
compare, for example, TikiWiki syntax:
*One
*Two
**Two-sub
There is no way to put other things in a list. It begins and ends when
the list items stop.
>
> I upgraded from CVS including your patch commited by arnaud. I installed
> it and run the provided example. Simply copied/paste the integrated
> help. Pehaps your <li> <ul> 's are better Xhtml, but they are _wrong_,
> adding in all case a first empty element and wrong numbering (the stack
> :) ? )
>
> > The best way to do this is going to be adding a stack based parser in
> > there somewhere. I would propose a "validate" step between the
> > "parse" and "render" steps that can add, remove, and rearrange tokens
> > as needed.
>
> Good catch, we need some structure's check before final rendering.
> There's no best way. I still see no utility in a "stack". We can do that
> using the existing structure and check the proper nesting.
>
> >
> > When you put it there, you can store the results, which is nice
> > because stack based parsing like what needs doing is rather compute
> > intensive.
>
Text_Wiki is only a text processor. It converts from one format to
another. If you give it bad input you're going to get bad output.
Text_Wiki itself knows nothing about validity of structures. It
*could* be possible to validate the input....possibly....but this
would be something that should be done as little as possible due to
the extra overhead.
If you want to see an interesting solution for fixing the XML-ness of
HTML, here's a snippet of code I wrote:
http://pear.reversefold.com/fixHtml.php
It makes sure that all XML tags are ended/started correctly and even
adds more tags to (hopefully) keep the original intent
(<b>bold<i>italic bold</b>italic<i> becomes <b>bold<i>italic
bold</i></b><i>italic<i> for example). It's not perfect, but it does
fix XML validity at least.
Now I understand the "need" for verifying syntax and output, but I can
see only one way that Text_Wiki can conceivably do this. First of all,
I'll restate that I think this should be *optional* and off by
default. Verification should only be done upon input and not on normal
output, otherwise you're going to be verifying the same thing over and
over again. (This is, of course, an application-level decision, which
is why I'm stressing the need for flexibility.) Second, this
verification should probably be done on the intermediate format. This
way the verification need not be specific to any input or output
syntax.
For this to work, we (the Text_Wiki devs) will have to alter some of
the "tokens" to be more standard. To verify the "XML-ness" of the
tokens we need to have the 'type' key for start and end tokens be the
same and have a key which is always 'start' or 'end' so that we can
easily match them *without* knowing specifics about the tokens. This
was we can use a simple stack-based approach for validation of
start/close tags. This should take care of validating start/end
tokens.
The harder part is validation of XHtml structures. I don't think that
this belongs in Text_Wiki at all. It's simply out of scope. Text_Wiki
doesn't really know anything about the input and output formats, it
only converts them (have I said this before?). We *could* add an
optional validator, say, for the different input syntaxes, but I have
no idea how this could be sanely implemented. IMHO this should be left
to the application to do. If you're that worried about XHtml
compliance (I try not to when I'm letting users edit things. They will
*never* all understand what XHtml is and why they should enter things
a certain way.) then the output should be run through tidy or some
other validation engine and rejected if it is incorrect. Or you could
allow tidy to fix it (make it XHtml) and cache *that* as the output.
--
Justin Patrin