Re: BBCodeParser (transition to Text_Wiki)
| From: | Justin Patrin | Date: | Fri, 14 Oct 2005 18:41:48 +0000 |
| Subject: | Re: BBCodeParser (transition to Text_Wiki) | ||
| References: | 1 2 3 4 5 6 | Groups: | php.pear.dev |
| Request: | Send a blank email to pear-dev+get-40196@lists.php.net to get a copy of this message | ||
Please reply to the thread and don't copy/paste things you want to
respond to. Respond in-line in the e-mail below. Having 2 copies of
the previous comments is annoying.
(responses below)
On 10/14/05, Seth Price <seth@pricepages.org> wrote:
> > Garbage In, Garbage Out! The XHtml renderer outputs XHtml if you give
> > it good input. If you're worried about it I suggest you run the output
> > through tidy (or similar) upon submission and reject it if the XML is
> > broken. It doesn't have to be validated every time you render it, only
> > when it's submitted. This leads me to believe it's outside the scope
> > of Text_Wiki.
> But it doesn't have to be like this. I'm aiming at supporting code
> written by people that don't know XML. I don't want to force people
> to learn proper XHTML before they submit formatted text with my site.
Agreed.
> If I can do it with BBCodeParser, then we can do it with Text_Wiki.
> With my current design, it is also only validated when submitted.
>
> > Same as above. I don't agree that this is an XSS attack. An XSS attack
> > injects unwanted code into yout page. One thing which *could* be
> > construed as an XSS attack would be if a parser allows unmatched tags
> > like [i] to be rendered without an ending [/i]. I don't *think* that
> > any of our parsers allow this but if they do, submit a bug report.
> One of the purposes for XHTML is that it can be rendered by an
> lightweight XML engine. If it doesn't match the DTD, then we have an
> error. Most web browsers are smarter than just displaying a rendering
> error message, but if someone can input text that forces the browser
> to correct the XHTML or error out, then that person has broken your
> page. Technically, it would be perfectly valid for the browser do
> display only a XML rendering error message.
True.
>
> > Yep, upon looking it seems that BBCode has a [list] [/list] format.
> > This is simply bad syntax on the part of BBCode. BBCode is the thing
> > allowing other "tags" within the [list] syntax, not Text_Wiki. If you
> > compare, for example, TikiWiki syntax:
> XHTML requires <ul> and </ul> tags, but I don't consider that "bad
> syntax on the part of" XHTML.
No, but you're missing my point. To write XHtml you have to understand
it. Writing bad XHtml causes an error in some situations. This is
simply how it is. Writing bad PHP also results in an error. That's the
point of having a syntax for things.
Now think about BBCode. What is its point? What is its value? One of
the major points is to simplify markup to be easily used by users.
Now, what actually happens with BBCode? You *still* have to understand
the basics of XML to know how to write BBCode correctly. This is a
failure of BBCode to do what it is supposed to do. Hence it's BBCode's
fault for not doing what it proposes to do.
What I'm saying is that if you want something for your users to use in
order to make it easier for them (and to keep from instriducing
errors) then you need to choose the right tool for the job. If BBCode
si structured such that it *can be written badly* then perhaps you
should consider using something else. Such as Dokuwiki or TIkiWiki
syntax (although they may also have some similar problems which should
be addressed). My point here is that we're trying to fix shortcomings
of the *underlying syntax* with hacks and hacks only complicate
things.
>
> > If you want to see an interesting solution for fixing the XML-ness of
> > HTML, here's a snippet of code I wrote:
> > http://pear.reversefold.com/fixHtml.php
> That's starting to look similar to the stack based engine in
> HTML_BBCodeParser ;)
>
> > Now I understand the "need" for verifying syntax and output, but I can
> > [snip]
> Yep, this is the same thought process that I have had, but I still
> want to work in a validation engine for XHTML structures somewhere.
> Part of me thinks that this should be built into the BBCode parser,
> because:
> [*] list item
>
> should be turned into "proper" BBCode:
> [list][li]list item[/li][/list]
Now you're talking about a whole new thing. You're saying that we
should be able to read the intent of the user from flawed markup. This
a large part of why web browsers are so f***ed up today. They allow
for bad markup and they read an intent into them rather than sticking
to a standard and throwing away things which don't comply.
Taking [*] which is not surrounded by [list] and magically adding
[list] is a magic step which doesn't really belong in a simple
translator (yes, Text_Wiki is a simple translator). You would have to
do such a thing in a preprocessor step since the List rules in
Text_Wiki are built to handle an entire list at once and not to see
[*] as an <li> all by itself. I suggest we leave such things as they
are. If a user sees it and wonders why they can go back and look at
the rules again and fix their markup.
>
> But then I'm sure that the various Wiki codes could benefit from
> proper XHTML structure validation also, so maybe it should be worked
> into the rendering engine.
As I tried to make clear earlier, if the syntax has been structured
correctly it shouldn't be possible to introduce bad output (except for
unbalanced tags, which are fairly easily fixed/checked for). And
again, I *do not* think that an XHtml validator is within the scope of
Text_Wiki.
> But then every page retrieval needs to re-
> correct the broken BBCode. And I'm not sure how Latex works, but I'm
> sure that it could use some structure validation. So it could go in
> either place.
Not sure what you mean. As I said before, I think that validation
should happen upon submission and if the input is broken you shouldn't
allow the user to save it. (Or allow them to save it as "plaintext"
which will be output as-is.) If you allow submission of bad code
you're going to have to jump through all sorts of hoops to make it
correct.
>
> Maybe we should implement a parser that can read XHTML, so then I can
> render it as XHTML with the structure validation, and then parse it
> back into intermediate format again for storage. It would look like
> this:
Blech. Yes, an XHtml parser is something we want, but this isn't IMHO
a good use for one. You're assuming here that we're not only going to
do structure validation but also code fixing on an intent level, which
I have made my stance clear on.
>
> -(messy BBCode)
> [*] list item
>
> 1. parse the BBCode and do tag matching in a stack
>
> -(intermediate format with improper structure)
> {li token}list item{/li token}
>
> 2. render as XHTML and do structure validation (most likely in
> another stack)
>
> -(XHTML 1.1 compliant)
> <ul><li>list item</li></ul>
>
> 3. parse the XHTML
>
> -(intermediate format with proper structure for XHTML)
> {list token}{li token}list item{/li token}{/list token}
>
> 4. store in database
>
> This would be a royal pain the the butt to implement, and would
> require both a stack based parser and a stack based render, but
> afterwards would have the best and most extensible Wiki/BBCode/XHTML
> parser/render available for use.
>
> Again, I don't know Latex, but the same method could be used to
> validate anything intended for Latex output with the addition of a
> Latex parser. You could render XHTML, parse XHTML, render Latex,
> parse Latex, and have intermediate format that is *guaranteed* valid
> in *both* Latex and XHTML. Oh... one can dream...
>
> Personally, I can't find a package out there that does a quality job
> rendering BBCode. I think the closest now is HTML_BBCodeParser after
> the patches that I made, but it still really isn't *that* good. I was
> just telling my aunt that getting a computer to understand a human is
> just as hard as getting a human to understand a computer. She seemed
> to understand. ;)
> ~Seth
>
> On Oct 14, 2005, at 11:58 AM, Justin Patrin wrote:
>
> > On 10/14/05, bertrand Gugger <bertrand@toggg.com> wrote:
> >
> >> Bonjour,
> >> Seth Price wrote:
> >>
> >>
> >
> > Responding more to Seth than bertrand...
> >
> >
> >>> Hey, I have a chance to look through the Text_Wiki code now.
> >>> Replying
> >>> to your comments:
> >>>
> >>>
> >>>> Actually, the whole structure was borked from the beginning, I see
> >>>> no point to reinvent a tree manager to parse BBCode, so it's
> >>>> vain to
> >>>> go ahead in this direction and make the code even more
> >>>> complicated.
> >>>> Text_Wiki 's treeing is implicit.
> >>>>
> >>>
> >>> I don't see how you can guarantee XHTML compliant output without
> >>> some
> >>> sort of stack based parser.
> >>>
> >>
> >> I don't guarantee anything, I use the Text_Wiki engine and its Xhtml
> >> renderer
> >>
> >
> > Garbage In, Garbage Out! The XHtml renderer outputs XHtml if you give
> > it good input. If you're worried about it I suggest you run the output
> > through tidy (or similar) upon submission and reject it if the XML is
> > broken. It doesn't have to be validated every time you render it, only
> > when it's submitted. This leads me to believe it's outside the scope
> > of Text_Wiki.
> >
> >
> >>
> >>
> >>> Simply put, if "[i][b]txt[/i][/b]" results in
> >>> "<i><b>txt</i></
> >>> b>" (as
> >>> it does now), then you can't claim that output will be XHTML.
> >>> Mismatched BBCode may even qualify as a XSS-style attack because it
> >>> breaks the validity of your output.
> >>>
> >
> > Same as above. I don't agree that this is an XSS attack. An XSS attack
> > injects unwanted code into yout page. One thing which *could* be
> > construed as an XSS attack would be if a parser allows unmatched tags
> > like [i] to be rendered without an ending [/i]. I don't *think* that
> > any of our parsers allow this but if they do, submit a bug report.
> >
> >
> >>>
> >>> You also need some way of guaranteeing that the only tag in <ul> is
> >>> <li> and other similar PEBKAC errors.
> >>>
> >
> > Text_Wiki *shouldn't* be putting anything inside lists. Then again,
> > I'm used to wiki markup which doesn't normally have a "start list" and
> > an "end list" tag.
> >
> > Yep, upon looking it seems that BBCode has a [list] [/list] format.
> > This is simply bad syntax on the part of BBCode. BBCode is the thing
> > allowing other "tags" within the [list] syntax, not Text_Wiki. If you
> > compare, for example, TikiWiki syntax:
> >
> > *One
> > *Two
> > **Two-sub
> >
> > There is no way to put other things in a list. It begins and ends when
> > the list items stop.
> >
> >
> >>
> >> I upgraded from CVS including your patch commited by arnaud. I
> >> installed
> >> it and run the provided example. Simply copied/paste the integrated
> >> help. Pehaps your <li> <ul> 's are better Xhtml, but they are
> >> _wrong_,
> >> adding in all case a first empty element and wrong numbering (the
> >> stack
> >> :) ? )
> >>
> >>
> >>> The best way to do this is going to be adding a stack based
> >>> parser in
> >>> there somewhere. I would propose a "validate" step between the
> >>> "parse" and "render" steps that can add, remove, and
> >>> rearrange
> >>> tokens
> >>> as needed.
> >>>
> >>
> >> Good catch, we need some structure's check before final rendering.
> >> There's no best way. I still see no utility in a "stack". We can
> >> do that
> >> using the existing structure and check the proper nesting.
> >>
> >>
> >>>
> >>> When you put it there, you can store the results, which is nice
> >>> because stack based parsing like what needs doing is rather compute
> >>> intensive.
> >>>
> >>
> >>
> >
> > Text_Wiki is only a text processor. It converts from one format to
> > another. If you give it bad input you're going to get bad output.
> > Text_Wiki itself knows nothing about validity of structures. It
> > *could* be possible to validate the input....possibly....but this
> > would be something that should be done as little as possible due to
> > the extra overhead.
> >
> > If you want to see an interesting solution for fixing the XML-ness of
> > HTML, here's a snippet of code I wrote:
> > http://pear.reversefold.com/fixHtml.php
> > It makes sure that all XML tags are ended/started correctly and even
> > adds more tags to (hopefully) keep the original intent
> > (<b>bold<i>italic bold</b>italic<i> becomes
> > <b>bold<i>italic
> > bold</i></b><i>italic<i> for example). It's not perfect, but
> > it does
> > fix XML validity at least.
> >
> > Now I understand the "need" for verifying syntax and output, but I can
> > see only one way that Text_Wiki can conceivably do this. First of all,
> > I'll restate that I think this should be *optional* and off by
> > default. Verification should only be done upon input and not on normal
> > output, otherwise you're going to be verifying the same thing over and
> > over again. (This is, of course, an application-level decision, which
> > is why I'm stressing the need for flexibility.) Second, this
> > verification should probably be done on the intermediate format. This
> > way the verification need not be specific to any input or output
> > syntax.
> >
> > For this to work, we (the Text_Wiki devs) will have to alter some of
> > the "tokens" to be more standard. To verify the "XML-ness" of the
> > tokens we need to have the 'type' key for start and end tokens be the
> > same and have a key which is always 'start' or 'end' so that we can
> > easily match them *without* knowing specifics about the tokens. This
> > was we can use a simple stack-based approach for validation of
> > start/close tags. This should take care of validating start/end
> > tokens.
> >
> > The harder part is validation of XHtml structures. I don't think that
> > this belongs in Text_Wiki at all. It's simply out of scope. Text_Wiki
> > doesn't really know anything about the input and output formats, it
> > only converts them (have I said this before?). We *could* add an
> > optional validator, say, for the different input syntaxes, but I have
> > no idea how this could be sanely implemented. IMHO this should be left
> > to the application to do. If you're that worried about XHtml
> > compliance (I try not to when I'm letting users edit things. They will
> > *never* all understand what XHtml is and why they should enter things
> > a certain way.) then the output should be run through tidy or some
> > other validation engine and rejected if it is incorrect. Or you could
> > allow tidy to fix it (make it XHtml) and cache *that* as the output.
> >
> > --
> > Justin Patrin
> >
> > --
> > PEAR Development Mailing List (http://pear.php.net/)
> > To unsubscribe, visit: http://www.php.net/unsub.php
> >
> >
> >
> >
>
>
--
Justin Patrin