New proposal: HTML_Miner

From: Date: Sun, 15 Jan 2006 14:07:20 +0000
Subject: New proposal: HTML_Miner
Groups: php.pear.dev 
Request: Send a blank email to pear-dev+get-40962@lists.php.net to get a copy of this message
Hello, I developed a PHP module for extracting blocks from a HTML file, I called it HTML_Miner. The class is very easy to use: you call the method parseBlock($nodeSpecs, $html) that returns the inner and outer HTML of the block you specified in $nodeSpecs. An example: parseBlock("body->table->table", file_get_contents("my.html")); Will return the flowwing array: array( 'outer' => '<table ...>CONTENT</table>', 'inner' => 'CONTENT' ) The class will return the correct block following the html tree structure. In this example the class will return the first table, nested in first table of the document's body. The grammar used for $nodeSpecs can handle little more complicated requests, for example you can have: parseBlock("body->table[3]->table", file_get_contents("my.html")); or maybe: parseBlock("body->table(class='myTable')->table(cellpadding='7')", file_get_contents("my.html")); The former will get the first table nested in the third table of the body, the second example will get the first table having cellpadding=7 in the first table of the document's body having class='myTable'. You can mix both of these options and create complex expressions for block parsing. Do you think that this class can be of any interest to the pear community? I was thinking about subscribing to pear and submitting the class properly commented and formatted. What do you suggest? Thank you. Regards, Niccolò Venturoli Rome, Italy

« previous php.pear.dev (#40962) next »