Parsing PDF files

From: Date: Fri, 31 Aug 2001 15:16:38 +0000
Subject: Parsing PDF files
Groups: php.windows 
Request: Send a blank email to php-windows+get-9155@lists.php.net to get a copy of this message
Wondering if anyone has tried to parse out a table of information from a PDF file? Is it a matter of opening the file, looping through its contents line-by-line looking for tags that demarcate table cell boundaries and extracting the relevant cell values? I figured if Google can index PDF content it must be possible to pull the content out something like one would an HTML file. Mostly wondering how much work might be involved and if there are any tricks that I should be aware of before I begin... Regards, Paul

« previous php.windows (#9155) next »