Re: Suggestion: HTML extractor

From: Date: Thu, 12 Jun 2003 06:42:03 +0000
Subject: Re: Suggestion: HTML extractor
References: 1  Groups: php.pear.dev 
Request: Send a blank email to pear-dev+get-17345@lists.php.net to get a copy of this message
It's probably a 5 minute hack on the current flexy tokenizer engine.. - just run the toElement call on the tokens.. otherwise have a look at XML_HTMLSax I did have a play writing bindings to HTML-Tidy, but other than being a bit faster than the Lexer in Flexy, it's a very big API to convert.. Regards Alan Ignatius Reilly wrote:
Hello, I do a lot of web content extraction in back-office production environments, with custom scripts that REGEX for HTML patterns. This approach is not fully satisfactory, because scripts break down with the slightest web page redesigns and are painful to maintain. Therefore I am embarking on another (hopefully better) approach: HTML-Tidy a web page's content into XHTML and parse this as an XML document to extract content. Key advantage: content will be simply defined by an XPath expression. I would be glad to later contribute a PEAR class for this purpose, if it is deemed useful. However, AFAIK, there is no PHP API for HTML-Tidy (written in C), so I couldn't find better that calling from the command line (PHP backtick operator). This is probably an unacceptable requirement for a PEAR class. So my question is: is there anyone working on a PHP extension (PECL ?) of HTML-Tidy or anything similar? I will be happy to receive any comment/ guidance. Thanks Ignatius J. Reilly
-- Can you help out? Need Consulting Services or Know of a Job? http://www.akbkhome.com

« previous php.pear.dev (#17345) next »