Re: Suggestion: HTML extractor

From: Date: Thu, 12 Jun 2003 06:47:35 +0000
Subject: Re: Suggestion: HTML extractor
References: 1  Groups: php.pear.dev 
Request: Send a blank email to pear-dev+get-17346@lists.php.net to get a copy of this message
On Thu, 12 Jun 2003, Ignatius Reilly wrote: > Hello, > > I do a lot of web content extraction in back-office production environments, > with custom scripts that REGEX for HTML patterns. > This approach is not fully satisfactory, because scripts break down with the > slightest web page redesigns and are painful to maintain. > > Therefore I am embarking on another (hopefully better) approach: HTML-Tidy a > web page's content into XHTML and parse this as an XML document to extract > content. Key advantage: content will be simply defined by an XPath > expression. with html_doc resp. html_doc_file you should be able to load a HTML as a DOM tree without the need of converting it to XHTML. Maybe that helps chregu > > I would be glad to later contribute a PEAR class for this purpose, if it is > deemed useful. > However, AFAIK, there is no PHP API for HTML-Tidy (written in C), so I > couldn't find better that calling from the command line (PHP backtick > operator). This is probably an unacceptable requirement for a PEAR class. > > So my question is: is there anyone working on a PHP extension (PECL ?) of > HTML-Tidy or anything similar? > > I will be happy to receive any comment/ guidance. > > Thanks > > Ignatius J. Reilly > > > -- nam...christian stocker adr...pflanzschulstr. 31, ch-8004 zurich pho...+41 43 317 9984 www...http://blog.bitflux.ch mob...+41 76 561 8860 ema...chregu@phant.ch wor...+41 1 240 5670 gpg...0x5CE1DECB

« previous php.pear.dev (#17346) next »