Garbage In, Garbage Out! The XHtml renderer outputs XHtml if you give
it good input. If you're worried about it I suggest you run the output
through tidy (or similar) upon submission and reject it if the XML is
broken. It doesn't have to be validated every time you render it, only
when it's submitted. This leads me to believe it's outside the scope
of Text_Wiki.
But it doesn't have to be like this. I'm aiming at supporting code written by people that don't know XML. I don't want to force people to learn proper XHTML before they submit formatted text with my site. If I can do it with BBCodeParser, then we can do it with Text_Wiki. With my current design, it is also only validated when submitted.
Same as above. I don't agree that this is an XSS attack. An XSS attack
injects unwanted code into yout page. One thing which *could* be
construed as an XSS attack would be if a parser allows unmatched tags
like [i] to be rendered without an ending [/i]. I don't *think* that
any of our parsers allow this but if they do, submit a bug report.
One of the purposes for XHTML is that it can be rendered by an lightweight XML engine. If it doesn't match the DTD, then we have an error. Most web browsers are smarter than just displaying a rendering error message, but if someone can input text that forces the browser to correct the XHTML or error out, then that person has broken your page. Technically, it would be perfectly valid for the browser do display only a XML rendering error message.
Yep, upon looking it seems that BBCode has a [list] [/list] format.
This is simply bad syntax on the part of BBCode. BBCode is the thing
allowing other "tags" within the [list] syntax, not Text_Wiki. If you
compare, for example, TikiWiki syntax:
XHTML requires <ul> and </ul> tags, but I don't consider that "bad syntax on the part of" XHTML.
If you want to see an interesting solution for fixing the XML-ness of
HTML, here's a snippet of code I wrote:
http://pear.reversefold.com/fixHtml.php
That's starting to look similar to the stack based engine in HTML_BBCodeParser ;)
Now I understand the "need" for verifying syntax and output, but I can
[snip]
Yep, this is the same thought process that I have had, but I still want to work in a validation engine for XHTML structures somewhere. Part of me thinks that this should be built into the BBCode parser, because:
[*] list item
should be turned into "proper" BBCode:
[list][li]list item[/li][/list]
But then I'm sure that the various Wiki codes could benefit from proper XHTML structure validation also, so maybe it should be worked into the rendering engine. But then every page retrieval needs to re-correct the broken BBCode. And I'm not sure how Latex works, but I'm sure that it could use some structure validation. So it could go in either place.
Maybe we should implement a parser that can read XHTML, so then I can render it as XHTML with the structure validation, and then parse it back into intermediate format again for storage. It would look like this:
-(messy BBCode)
[*] list item
1. parse the BBCode and do tag matching in a stack
-(intermediate format with improper structure)
{li token}list item{/li token}
2. render as XHTML and do structure validation (most likely in another stack)
-(XHTML 1.1 compliant)
<ul><li>list item</li></ul>
3. parse the XHTML
-(intermediate format with proper structure for XHTML)
{list token}{li token}list item{/li token}{/list token}
4. store in database
This would be a royal pain the the butt to implement, and would require both a stack based parser and a stack based render, but afterwards would have the best and most extensible Wiki/BBCode/XHTML parser/render available for use.
Again, I don't know Latex, but the same method could be used to validate anything intended for Latex output with the addition of a Latex parser. You could render XHTML, parse XHTML, render Latex, parse Latex, and have intermediate format that is *guaranteed* valid in *both* Latex and XHTML. Oh... one can dream...
Personally, I can't find a package out there that does a quality job rendering BBCode. I think the closest now is HTML_BBCodeParser after the patches that I made, but it still really isn't *that* good. I was just telling my aunt that getting a computer to understand a human is just as hard as getting a human to understand a computer. She seemed to understand. ;)
~Seth
On Oct 14, 2005, at 11:58 AM, Justin Patrin wrote:
On 10/14/05, bertrand Gugger <bertrand@toggg.com> wrote:
Bonjour,
Seth Price wrote:
Responding more to Seth than bertrand...
Hey, I have a chance to look through the Text_Wiki code now. Replying
to your comments:
Actually, the whole structure was borked from the beginning, I see
no point to reinvent a tree manager to parse BBCode, so it's vain to
go ahead in this direction and make the code even more complicated.
Text_Wiki 's treeing is implicit.
I don't see how you can guarantee XHTML compliant output without some
sort of stack based parser.
I don't guarantee anything, I use the Text_Wiki engine and its Xhtml
renderer
Garbage In, Garbage Out! The XHtml renderer outputs XHtml if you give
it good input. If you're worried about it I suggest you run the output
through tidy (or similar) upon submission and reject it if the XML is
broken. It doesn't have to be validated every time you render it, only
when it's submitted. This leads me to believe it's outside the scope
of Text_Wiki.
Simply put, if "[i][b]txt[/i][/b]" results in "<i><b>txt</i></b>" (as
it does now), then you can't claim that output will be XHTML.
Mismatched BBCode may even qualify as a XSS-style attack because it
breaks the validity of your output.
Same as above. I don't agree that this is an XSS attack. An XSS attack
injects unwanted code into yout page. One thing which *could* be
construed as an XSS attack would be if a parser allows unmatched tags
like [i] to be rendered without an ending [/i]. I don't *think* that
any of our parsers allow this but if they do, submit a bug report.
You also need some way of guaranteeing that the only tag in <ul> is
<li> and other similar PEBKAC errors.
Text_Wiki *shouldn't* be putting anything inside lists. Then again,
I'm used to wiki markup which doesn't normally have a "start list" and
an "end list" tag.
Yep, upon looking it seems that BBCode has a [list] [/list] format.
This is simply bad syntax on the part of BBCode. BBCode is the thing
allowing other "tags" within the [list] syntax, not Text_Wiki. If you
compare, for example, TikiWiki syntax:
*One
*Two
**Two-sub
There is no way to put other things in a list. It begins and ends when
the list items stop.
I upgraded from CVS including your patch commited by arnaud. I installed
it and run the provided example. Simply copied/paste the integrated
help. Pehaps your <li> <ul> 's are better Xhtml, but they are _wrong_,
adding in all case a first empty element and wrong numbering (the stack
:) ? )
The best way to do this is going to be adding a stack based parser in
there somewhere. I would propose a "validate" step between the
"parse" and "render" steps that can add, remove, and rearrange tokens
as needed.
Good catch, we need some structure's check before final rendering.
There's no best way. I still see no utility in a "stack". We can do that
using the existing structure and check the proper nesting.
When you put it there, you can store the results, which is nice
because stack based parsing like what needs doing is rather compute
intensive.
Text_Wiki is only a text processor. It converts from one format to
another. If you give it bad input you're going to get bad output.
Text_Wiki itself knows nothing about validity of structures. It
*could* be possible to validate the input....possibly....but this
would be something that should be done as little as possible due to
the extra overhead.
If you want to see an interesting solution for fixing the XML-ness of
HTML, here's a snippet of code I wrote:
http://pear.reversefold.com/fixHtml.php
It makes sure that all XML tags are ended/started correctly and even
adds more tags to (hopefully) keep the original intent
(<b>bold<i>italic bold</b>italic<i> becomes <b>bold<i>italic
bold</i></b><i>italic<i> for example). It's not perfect, but it does
fix XML validity at least.
Now I understand the "need" for verifying syntax and output, but I can
see only one way that Text_Wiki can conceivably do this. First of all,
I'll restate that I think this should be *optional* and off by
default. Verification should only be done upon input and not on normal
output, otherwise you're going to be verifying the same thing over and
over again. (This is, of course, an application-level decision, which
is why I'm stressing the need for flexibility.) Second, this
verification should probably be done on the intermediate format. This
way the verification need not be specific to any input or output
syntax.
For this to work, we (the Text_Wiki devs) will have to alter some of
the "tokens" to be more standard. To verify the "XML-ness" of the
tokens we need to have the 'type' key for start and end tokens be the
same and have a key which is always 'start' or 'end' so that we can
easily match them *without* knowing specifics about the tokens. This
was we can use a simple stack-based approach for validation of
start/close tags. This should take care of validating start/end
tokens.
The harder part is validation of XHtml structures. I don't think that
this belongs in Text_Wiki at all. It's simply out of scope. Text_Wiki
doesn't really know anything about the input and output formats, it
only converts them (have I said this before?). We *could* add an
optional validator, say, for the different input syntaxes, but I have
no idea how this could be sanely implemented. IMHO this should be left
to the application to do. If you're that worried about XHtml
compliance (I try not to when I'm letting users edit things. They will
*never* all understand what XHtml is and why they should enter things
a certain way.) then the output should be run through tidy or some
other validation engine and rejected if it is incorrect. Or you could
allow tidy to fix it (make it XHtml) and cache *that* as the output.
--
Justin Patrin
--
PEAR Development Mailing List (http://pear.php.net/)
To unsubscribe, visit: http://www.php.net/unsub.php