Re: [PEPr] Comment on PDF::PDF_Parser

From: Date: Mon, 21 Jun 2004 19:42:02 +0000
Subject: Re: [PEPr] Comment on PDF::PDF_Parser
References: 1  Groups: php.pear.dev 
Request: Send a blank email to pear-dev+get-31054@lists.php.net to get a copy of this message
Hello, Am 21.06.2004 18:11 schrieb PEPr: > It'd be great if the object model/potential writing capabilities > could be integrated with the existing PDF package, so that you could > take a .pdf file, parse it into objects, work on it with the PDF > class, and write it out again. Any chance of that integration > happening? This integration can be done, but it depends also on Marko Djukic, because he maintains the File_PDF package, and I want to avoid a duplication of the packages functionality. Thats the reason for the very simple implementation of the Renderer in PDF_PageExtractor. On the other hand the current PDF_Parser can't process content streams. He is designed for Object (PDF_Lexer) and File Structure (PDF_Parser) parsing. The package PDF_PageExtractor implements a very basic Document Structure parser, which doesn't support all available features (i.e. Outlines or Page Labels), which i will implement in any case. Additional to this you need a content stream parser, if you want to edit text directly. See [1] for more in information. The content stream parser will be the hardest part to implement, you have to learn much about color theory, color spaces, fonts, etc. [2] For a simple text extractor a small parser can be created, but thats not on top of my todo list. I would like to know, if the choosen names were ok? Or should the whole packages integrated under the File_* hierachy? My favourite structure would look like this: PDF +-Objects.php +-Objects (dir) | +-bool.php | +-collection.php (dictionary, reference, stream, array) | +-float.php | +-integer.php | +-null.php | +-text (string, name, comment, keyword) | +-Structures.php +-Structures (dir) | +-pages.php | +-page.php | +-font.php | +-annotation.php | +-outlines.php | +-thumbnails.php | +-(tbd).php | +-Parser.php +-Parser (dir) | +-objects.php (current PDF_Lexer) | +-file.php (current PDF_Parser) | +-document.php (basics in PDF_PageExtractor) | | +-document (dir) | | +-pages.php | | +-outlines.php | | +-article.php | | +-forms.php | | +-(tbd).php | +-content.php | +-Buffer.php | +-Buffer (dir) | +-file.php | +-string.php | +-Renderer.php (very simple atm) +-Renderer (dir) | +-(different for objects and structures) | +-Utilities (dir) +-PageExtractor.php +-ImageExtractor.php +-TextExtractor.php +-Merger.php +-Template.php +-Creator.php (simple API for creating documents) +-(what you imagine).php As you can see i would prefer a flat hierachy. For example you can build a document structure with the objects and structures, which is outputed throu the renderer. A clever programmer can write a html2pdf converter by converting html tags to the pdf objects and then outputing thru the renderer. If here is a approval for this structure i would change my proposal to reflect this. With the current state of the package (including PDF_PageExtractor) i think, following utilities would be possible: - PageExtractor (see proposal) - ImageExtractor (only .jpg, .bmp and .png) - a simple Merger - a very simple Template, where white spaces are overwritten, for this you have to know the exact position of the white boxes (this can be implemented by using the File_PDF package, with some on the output algorithm) I hope nobody fall asleep ;) regards Stefan [1] Figure 3.1, p. 24 PDF Reference fourth edition [2] Chapter 4 to 7 of PDF Reference fourth edition

« previous php.pear.dev (#31054) next »