Re: [PEPr] Comment on PDF::PDF_Parser
| From: | Stefan Ohrmann | Date: | Mon, 21 Jun 2004 19:42:02 +0000 |
| Subject: | Re: [PEPr] Comment on PDF::PDF_Parser | ||
| References: | 1 | Groups: | php.pear.dev |
| Request: | Send a blank email to pear-dev+get-31054@lists.php.net to get a copy of this message | ||
Hello,
Am 21.06.2004 18:11 schrieb PEPr:
> It'd be great if the object model/potential writing capabilities
> could be integrated with the existing PDF package, so that you could
> take a .pdf file, parse it into objects, work on it with the PDF
> class, and write it out again. Any chance of that integration
> happening?
This integration can be done, but it depends also on Marko Djukic,
because he maintains the File_PDF package, and I want to avoid a
duplication of the packages functionality. Thats the reason for the very
simple implementation of the Renderer in PDF_PageExtractor.
On the other hand the current PDF_Parser can't process content streams.
He is designed for Object (PDF_Lexer) and File Structure (PDF_Parser)
parsing. The package PDF_PageExtractor implements a very basic Document
Structure parser, which doesn't support all available features (i.e.
Outlines or Page Labels), which i will implement in any case. Additional
to this you need a content stream parser, if you want to edit text
directly. See [1] for more in information. The content stream parser
will be the hardest part to implement, you have to learn much about
color theory, color spaces, fonts, etc. [2] For a simple text extractor
a small parser can be created, but thats not on top of my todo list.
I would like to know, if the choosen names were ok? Or should the whole
packages integrated under the File_* hierachy? My favourite structure
would look like this:
PDF
+-Objects.php
+-Objects (dir)
| +-bool.php
| +-collection.php (dictionary, reference, stream, array)
| +-float.php
| +-integer.php
| +-null.php
| +-text (string, name, comment, keyword)
|
+-Structures.php
+-Structures (dir)
| +-pages.php
| +-page.php
| +-font.php
| +-annotation.php
| +-outlines.php
| +-thumbnails.php
| +-(tbd).php
|
+-Parser.php
+-Parser (dir)
| +-objects.php (current PDF_Lexer)
| +-file.php (current PDF_Parser)
| +-document.php (basics in PDF_PageExtractor)
| | +-document (dir)
| | +-pages.php
| | +-outlines.php
| | +-article.php
| | +-forms.php
| | +-(tbd).php
| +-content.php
| +-Buffer.php
| +-Buffer (dir)
| +-file.php
| +-string.php
|
+-Renderer.php (very simple atm)
+-Renderer (dir)
| +-(different for objects and structures)
|
+-Utilities (dir)
+-PageExtractor.php
+-ImageExtractor.php
+-TextExtractor.php
+-Merger.php
+-Template.php
+-Creator.php (simple API for creating documents)
+-(what you imagine).php
As you can see i would prefer a flat hierachy. For example you can build
a document structure with the objects and structures, which is outputed
throu the renderer. A clever programmer can write a html2pdf converter
by converting html tags to the pdf objects and then outputing thru the
renderer. If here is a approval for this structure i would change my
proposal to reflect this.
With the current state of the package (including PDF_PageExtractor) i
think, following utilities would be possible:
- PageExtractor (see proposal)
- ImageExtractor (only .jpg, .bmp and .png)
- a simple Merger
- a very simple Template, where white spaces are overwritten, for this
you have to know the exact position of the white boxes (this can be
implemented by using the File_PDF package, with some on the output
algorithm)
I hope nobody fall asleep ;)
regards
Stefan
[1] Figure 3.1, p. 24 PDF Reference fourth edition
[2] Chapter 4 to 7 of PDF Reference fourth edition