Suggestion: HTML extractor
| From: | Ignatius Reilly | Date: | Thu, 12 Jun 2003 06:09:42 +0000 |
| Subject: | Suggestion: HTML extractor | ||
| Groups: | php.pear.dev | ||
| Request: | Send a blank email to pear-dev+get-17344@lists.php.net to get a copy of this message | ||
Hello,
I do a lot of web content extraction in back-office production environments,
with custom scripts that REGEX for HTML patterns.
This approach is not fully satisfactory, because scripts break down with the
slightest web page redesigns and are painful to maintain.
Therefore I am embarking on another (hopefully better) approach: HTML-Tidy a
web page's content into XHTML and parse this as an XML document to extract
content. Key advantage: content will be simply defined by an XPath
expression.
I would be glad to later contribute a PEAR class for this purpose, if it is
deemed useful.
However, AFAIK, there is no PHP API for HTML-Tidy (written in C), so I
couldn't find better that calling from the command line (PHP backtick
operator). This is probably an unacceptable requirement for a PEAR class.
So my question is: is there anyone working on a PHP extension (PECL ?) of
HTML-Tidy or anything similar?
I will be happy to receive any comment/ guidance.
Thanks
Ignatius J. Reilly