Re: Suggestion: HTML extractor
| From: | Christian Stocker | Date: | Thu, 12 Jun 2003 06:47:35 +0000 |
| Subject: | Re: Suggestion: HTML extractor | ||
| References: | 1 | Groups: | php.pear.dev |
| Request: | Send a blank email to pear-dev+get-17346@lists.php.net to get a copy of this message | ||
On Thu, 12 Jun 2003, Ignatius Reilly wrote:
> Hello,
>
> I do a lot of web content extraction in back-office production environments,
> with custom scripts that REGEX for HTML patterns.
> This approach is not fully satisfactory, because scripts break down with the
> slightest web page redesigns and are painful to maintain.
>
> Therefore I am embarking on another (hopefully better) approach: HTML-Tidy a
> web page's content into XHTML and parse this as an XML document to extract
> content. Key advantage: content will be simply defined by an XPath
> expression.
with html_doc resp. html_doc_file you should be able to load a HTML as a
DOM tree without the need of converting it to XHTML.
Maybe that helps
chregu
>
> I would be glad to later contribute a PEAR class for this purpose, if it is
> deemed useful.
> However, AFAIK, there is no PHP API for HTML-Tidy (written in C), so I
> couldn't find better that calling from the command line (PHP backtick
> operator). This is probably an unacceptable requirement for a PEAR class.
>
> So my question is: is there anyone working on a PHP extension (PECL ?) of
> HTML-Tidy or anything similar?
>
> I will be happy to receive any comment/ guidance.
>
> Thanks
>
> Ignatius J. Reilly
>
>
>
--
nam...christian stocker adr...pflanzschulstr. 31, ch-8004 zurich
pho...+41 43 317 9984 www...http://blog.bitflux.ch
mob...+41 76 561 8860 ema...chregu@phant.ch
wor...+41 1 240 5670 gpg...0x5CE1DECB