Re: BBCodeParser (transition to Text_Wiki)

From: Date: Fri, 14 Oct 2005 19:37:04 +0000
Subject: Re: BBCodeParser (transition to Text_Wiki)
References: 1 2 3 4 5 6 7  Groups: php.pear.dev 
Request: Send a blank email to pear-dev+get-40199@lists.php.net to get a copy of this message
On Oct 14, 2005, at 1:41 PM, Justin Patrin wrote:
Please reply to the thread and don't copy/paste things you want to respond to. Respond in-line in the e-mail below. Having 2 copies of the previous comments is annoying.
I find it more difficult to read this way, but if you prefer it...
(responses below) On 10/14/05, Seth Price <seth@pricepages.org> wrote:
Garbage In, Garbage Out! The XHtml renderer outputs XHtml if you give it good input. If you're worried about it I suggest you run the output through tidy (or similar) upon submission and reject it if the XML is broken. It doesn't have to be validated every time you render it, only when it's submitted. This leads me to believe it's outside the scope of Text_Wiki.
But it doesn't have to be like this. I'm aiming at supporting code written by people that don't know XML. I don't want to force people to learn proper XHTML before they submit formatted text with my site.
Agreed.
If I can do it with BBCodeParser, then we can do it with Text_Wiki. With my current design, it is also only validated when submitted.
Same as above. I don't agree that this is an XSS attack. An XSS attack injects unwanted code into yout page. One thing which *could* be construed as an XSS attack would be if a parser allows unmatched tags like [i] to be rendered without an ending [/i]. I don't *think* that any of our parsers allow this but if they do, submit a bug report.
One of the purposes for XHTML is that it can be rendered by an lightweight XML engine. If it doesn't match the DTD, then we have an error. Most web browsers are smarter than just displaying a rendering error message, but if someone can input text that forces the browser to correct the XHTML or error out, then that person has broken your page. Technically, it would be perfectly valid for the browser do display only a XML rendering error message.
True.
Yep, upon looking it seems that BBCode has a [list] [/list] format. This is simply bad syntax on the part of BBCode. BBCode is the thing allowing other "tags" within the [list] syntax, not Text_Wiki. If you compare, for example, TikiWiki syntax:
XHTML requires <ul> and </ul> tags, but I don't consider that "bad syntax on the part of" XHTML.
No, but you're missing my point. To write XHtml you have to understand it. Writing bad XHtml causes an error in some situations. This is simply how it is. Writing bad PHP also results in an error. That's the point of having a syntax for things. Now think about BBCode. What is its point? What is its value? One of the major points is to simplify markup to be easily used by users. Now, what actually happens with BBCode? You *still* have to understand the basics of XML to know how to write BBCode correctly. This is a failure of BBCode to do what it is supposed to do. Hence it's BBCode's fault for not doing what it proposes to do. What I'm saying is that if you want something for your users to use in order to make it easier for them (and to keep from instriducing errors) then you need to choose the right tool for the job. If BBCode si structured such that it *can be written badly* then perhaps you should consider using something else. Such as Dokuwiki or TIkiWiki syntax (although they may also have some similar problems which should be addressed). My point here is that we're trying to fix shortcomings of the *underlying syntax* with hacks and hacks only complicate things.
It may be the fault of BBCode that it isn't "idiot proof" enough, but as the saying goes, if you make something more idiot proof, you will find a better idiot. If we are able to correct what problems we can, then we should do so.
If you want to see an interesting solution for fixing the XML-ness of HTML, here's a snippet of code I wrote: http://pear.reversefold.com/fixHtml.php
That's starting to look similar to the stack based engine in HTML_BBCodeParser ;)
Now I understand the "need" for verifying syntax and output, but I can [snip]
Yep, this is the same thought process that I have had, but I still want to work in a validation engine for XHTML structures somewhere. Part of me thinks that this should be built into the BBCode parser, because: [*] list item should be turned into "proper" BBCode: [list][li]list item[/li][/list]
Now you're talking about a whole new thing. You're saying that we should be able to read the intent of the user from flawed markup. This a large part of why web browsers are so f***ed up today. They allow for bad markup and they read an intent into them rather than sticking to a standard and throwing away things which don't comply.
Which is why I think that the logic for correcting bad syntax should be here, instead of depending on the web browser. The other choice that we have is simply not accepting the user input, but that isn't very user friendly.
Taking [*] which is not surrounded by [list] and magically adding [list] is a magic step which doesn't really belong in a simple translator (yes, Text_Wiki is a simple translator). You would have to do such a thing in a preprocessor step since the List rules in Text_Wiki are built to handle an entire list at once and not to see [*] as an <li> all by itself. I suggest we leave such things as they are. If a user sees it and wonders why they can go back and look at the rules again and fix their markup.
"[*] item" is easily matched with a regular expression. List item starts with a [*] and ends with a \n which gives you the [li]...[/li] tags. XHTML structure dictates that all <li> tags be enclosed by some sort of list, so HTML_BBCodeParser adds the <ul> list tags by default.
But then I'm sure that the various Wiki codes could benefit from proper XHTML structure validation also, so maybe it should be worked into the rendering engine.
As I tried to make clear earlier, if the syntax has been structured correctly it shouldn't be possible to introduce bad output (except for unbalanced tags, which are fairly easily fixed/checked for). And again, I *do not* think that an XHtml validator is within the scope of Text_Wiki.
Where is the best place for it? I am detecting that I just don't trust my users as much as you trust yours.
But then every page retrieval needs to re- correct the broken BBCode. And I'm not sure how Latex works, but I'm sure that it could use some structure validation. So it could go in either place.
Not sure what you mean. As I said before, I think that validation should happen upon submission and if the input is broken you shouldn't allow the user to save it. (Or allow them to save it as "plaintext" which will be output as-is.) If you allow submission of bad code you're going to have to jump through all sorts of hoops to make it correct.
Currently, we don't even have the ability to detect bad code in Text_Wiki. I guess the best way to do this is libtidy at some point.
Maybe we should implement a parser that can read XHTML, so then I can render it as XHTML with the structure validation, and then parse it back into intermediate format again for storage. It would look like this:
Blech. Yes, an XHtml parser is something we want, but this isn't IMHO a good use for one. You're assuming here that we're not only going to do structure validation but also code fixing on an intent level, which I have made my stance clear on.
It's not pretty, but it works. I want to fix code on an intent level, you don't, I'm ok with that. I'll use tidy_repair_string().
-(messy BBCode) [*] list item 1. parse the BBCode and do tag matching in a stack -(intermediate format with improper structure) {li token}list item{/li token} 2. render as XHTML and do structure validation (most likely in another stack) -(XHTML 1.1 compliant) <ul><li>list item</li></ul> 3. parse the XHTML -(intermediate format with proper structure for XHTML) {list token}{li token}list item{/li token}{/list token} 4. store in database This would be a royal pain the the butt to implement, and would require both a stack based parser and a stack based render, but afterwards would have the best and most extensible Wiki/BBCode/XHTML parser/render available for use. Again, I don't know Latex, but the same method could be used to validate anything intended for Latex output with the addition of a Latex parser. You could render XHTML, parse XHTML, render Latex, parse Latex, and have intermediate format that is *guaranteed* valid in *both* Latex and XHTML. Oh... one can dream... Personally, I can't find a package out there that does a quality job rendering BBCode. I think the closest now is HTML_BBCodeParser after the patches that I made, but it still really isn't *that* good. I was just telling my aunt that getting a computer to understand a human is just as hard as getting a human to understand a computer. She seemed to understand. ;) ~Seth


« previous php.pear.dev (#40199) next »