Re: Parsing PDF files

From: Date: Fri, 31 Aug 2001 16:56:18 +0000
Subject: Re: Parsing PDF files
References: 1  Groups: php.windows 
Request: Send a blank email to php-windows+get-9157@lists.php.net to get a copy of this message
I think that you can extract pretty easily the header, like: Subject, Creator, Author etc... But extracting values in a table may not that be so easy as the objects creation in the file are dependent on the file history and in addition the pdf file may be in a binary form. Alain On Fri, Aug 31, 2001 at 12:16:38PM -0300, Paul Meagher wrote: > Wondering if anyone has tried to parse out a table of information from a > PDF file? > > Is it a matter of opening the file, looping through its contents > line-by-line looking for tags that demarcate table cell boundaries and > extracting the relevant cell values? > > I figured if Google can index PDF content it must be possible to pull the > content out something like one would an HTML file. > > Mostly wondering how much work might be involved and if there are any > tricks that I should be aware of before I begin... > > Regards, > Paul > > > > > > > > > -- > PHP Windows Mailing List (http://www.php.net/) > To unsubscribe, e-mail: php-windows-unsubscribe@lists.php.net > For additional commands, e-mail: php-windows-help@lists.php.net > To contact the list administrators, e-mail: php-list-admin@lists.php.net

« previous php.windows (#9157) next »