Re: Text_Markup , some necessary evolution of Text_Wiki ?
| From: | Seth Price | Date: | Mon, 23 Jan 2006 16:50:00 +0000 |
| Subject: | Re: Text_Markup , some necessary evolution of Text_Wiki ? | ||
| References: | 1 2 3 | Groups: | php.pear.dev |
| Request: | Send a blank email to pear-dev+get-41052@lists.php.net to get a copy of this message | ||
You guys are talking about a version 2 of Text_Wiki. I have already written a version 2 of HTML_BBCodeParser, and have been thinking about version 3. I think we have a bit that we can learn from each other's packages. It sounds like you are mis-reading my words, and assuming that it my HTML_BBCodeParser is only applicable to BBCode because that is the package name. This is not so.
* It is not tied to BBCode, most of the code is only tied to stack based parsing. *
The only code that is BBCode specific in the core is the parser. And the parser can easily be changed to parse any tag-based code. (And because it can parse all tags with only one regex, it is much faster than the old parser or the Text_Wiki parser.)
A good object oriented design in Text_Wiki will allow for both Wiki style tags, and BBCode/HTML style ones. Therefore, I think my parser would be a good place to start for the BBCode/HTML tags, because of its efficiency and speed.
* My version of HTML_BBCodeParser is also not tied to HTML. *
Just a few days ago I finished up an ASCII renderer for it. It is more advanced than other ASCII renderers because it implements a basic box model, and therefore can handle proper indenting of lists. It also does line wrapping, to fit email widths. And it would not be possible without a stack based validator.
You might realize at this point that HTML_BBCodeParser has outgrown almost all jargon in its name, so that is why I was considering the name Text_Markup for it. I think I have code that you could use.
Just a few days ago, I started designing what could amount to a version 3 of HTML_BBCodeParser. Here are my rough notes about requirements and design. It is similar to what Paul linked to. The first half was written in a post to the QA list. The second half was written simply to remind myself the design I was thinking of when I have time to do more work on it.
From email Jan, 19th '06:
Examples of other things that I may like changed that would require possible BC breaks [in my version of HTML_BBCodeParser]:
- Stack based renderer
- This would allow easy use of things like Text_Highlighter in the [code] tag. It is pretty, but I don't see a good way to use the Highlighter as is.
- But the best implementation of the new renderer would require more filter changes.
- Why are all filters extended from BBCodeParser itself? It would make more sense to simply use a BBCodeParser_Filter class. This has led to odd bugs like #5844.
- Some of the tags behave unexpectedly, at least if you are used to vBulletin, phpBB, and/or InfoPop style tags (example: [quote] tags). Now would be a good time to change that functionality.
- Should newlines automatically be replaced with <br /> by default? That is what I would expect. (But they aren't.)
- The parse() method seems mis-named. Not only does it parse, but it verifies and renders too. We can keep qparse(), but I think that the other methods should be sorted out to allow for more control. More control would be useful when storing already-parsed-and-verified-but-not-rendered tokens in database for performance reasons.
- The current code for adding and removing filters on-the-fly works, but is kind of a hack to get around odd design stuff. I could rework that.
- The filter's _definedTags format seems kind of dumb in places. I could think of some more appropriate flags for tags.
- While I'm at it, I should take a look at Text_Wiki again. I remember there were some nice things in there from when I was looking at that code...
More stuff Jan, 20th '06 [design ideas]:
- Stack based parser will do better with tags like [raw] and [pre] also
- Maybe a object approch similar to Text_Wiki, but stack based stuff, like HTML_BBCodeParser2
- Parsing:
- Text_Markup calls Text_Markup_Parse_BBCode (subclass of Text_Markup_Parse)
- Text_Markup_Parse_BBCode calls all needed objects, like Text_Markup_Parse_BBCode_Align. It is better this way because then Text_Markup_Parse_BBCode can parse out all BBCode tags with one regex.
- Other markup languages without tags can then act more like the Text_Wiki parser and regex one token at a time.
- Validating:
- Stack based validating through the renderer(s).
- Call validate on whichever renderer you want to use, and it should have the nessesary information to validate the tree. Keep the validator from HTML_BBCodeParser
- If you want to render to more than one renderer, then call validate on each one of them.
- Text_Markup_Render_XHTML::validate() subclass of Text_Markup_Render::validate(), so validate can be overridden if needed. I can't think of why that would be needed though.
- Rendering:
- Text_Markup_Render_XHTML::render() subclass of Text_Markup_Render::validate()
- We can then use a stack based renderer, and do neat things...
- Output:
- We can export anything rendered, but but for performance reasons we might want to store validated text in the database.
- Thus, we may want a serialized array as output.
- But we should also include the version number so we can re-validate if there is a lower version number.
- To re-validate, we will also need to store enabled filter information and other options.
~Seth
On Jan 23, 2006, at 10:04 AM, bertrand Gugger wrote:
Seth Price wrote:It sounds like you are agreeing with my emails to the PEAR QA list about my version of HTML_BBCodeParser.No, I replied as I could , once. You finally persuded me this HTML_BBCodeParser is definitively deprecated or at least very oriented.Here are links to copies of the relevant emails.I follow all mails, they are nice archived now. Please, remind this thread is not specifical to BBCode. (I did not quote what is planned for this peculiar parser) Don't take my rudeness bad, I'm more about efficiency, possibly short minded. For the note, I have really some 100K sources I built only for the art of it , I don't care they are in trash, so wouldn't I for your code. I talk that way only in some hope of more cooperation. à+ --toggg