Re: Openoffice-to-docbook converter
| From: | Alan Knowles | Date: | Fri, 04 Apr 2003 07:12:53 +0000 |
| Subject: | Re: Openoffice-to-docbook converter | ||
| References: | 1 | Groups: | php.pear.dev |
| Request: | Send a blank email to pear-dev+get-14871@lists.php.net to get a copy of this message | ||
Thought I'd cc top pear dev just incase anyone else want to add anything.
xml.openoffice.org/xmerge/docbook/
(note openoffice's site is not the most stable around. = probably on constant attack from MS's script kiddies)
contains quite a bit of information about this - including a prelimary java extension to openoffice (which just crashed and died a horible death while trying to install itself - like most java programs :)
however the site does contain a good reference on the subject - namely:
a) a sample openoffice document in 'docbook format'
b) the docbook source it came from.
The background info is that I was working on a openoffice.org (get the name right :) to HTML convertor - that basically splits the documents based on Heading styles into pages, and I've started looking at extending that tool to deal with OOo-> <- docbook.
I started looking at the docbook -> OOo a few evenings ago.
So far:
- you point it at the package/en/database/db-dataobject.xml file
- it merges ../../../chapters.ent into the file (so the &xxxx; dont produce errors.)
- when it hits the &xxx; entities, it uses an extended cdata method to do a recursive parse (still looking at the best method for this - whether it's worth pre-parsing and creating the base file..)
- it then should go through and just create a content.xml file from a template or something.
- need to work out how to build zip files in php (as the built in extension only reads them)
For the OOo->docbook
So far: - not started - details below detail how the html one works
- based on the OOo->html converter
- opens the content.xml file from the OOo zip file
- sends it through a modified XML_Transformer (eg. extended and quite a few methods overwritten)
- has various namespace handlers which deal with OOo's namespaces..
- add <!-- SECTION xxx --> to the source as it goes along to indicate file breaks.
- final pass just preg_splits's on the <!-- SECTION tags, and puts them into files.
This is the current oo->html code http://devel.akbkhome.com/OO2html.html
usage is pretty simple
|$t = new XML_Transformer_OO2Html;
print_r($t->getIndex('/tmp/Xipe.sxw'));
||print_r($t->getPage('/tmp/Xipe.sxw','1'));
...
||
|
Regards
Alan
Ant-1 wrote:
Hi ! I had your email from Wolfram, who told me you are working on something that could end being close to an openoffice-to-docbook converter. I'm a member of Pear-doc, and we are all really looking forward to a tool like this. So what is exactly that you are planning and when do you think you could show an alpha ? regards, Antoine Pouch-- Can you help out? Need Consulting Services or Know of a Job? http://www.akbkhome.com