Looks like the same problem:
http://marc.theaimsgroup.com/?l=pear-dev&m=108739601312246&w=2
You need to use current CVS of PEAR.php, where the leak is fixed.
Thanks for advice, I tried it, but unfortunately it didn't help. To be sure I also tried to use XML_HTMLSax3, which doesn't extend from PEAR. I adjusted HTMLtoXHTML.php example from XML_HTMLSax3 package to check it. I parsed my own HTML file with size 154kB (example.html)
Excerpt from HTMLtoXHTML.php:
...
for ($i=0;$i<5;$i++) {
$memory = array();
$memory[] = memory_get_usage();
// Get the HTML file
$doc = file_get_contents('example.html');
$memory[] = memory_get_usage();
// Instantiate the handler
$handler=& new HTMLtoXHTMLHandler();
// Instantiate the parser
$parser=& new XML_HTMLSax3();
$memory[] = memory_get_usage();
// Register the handler with the parser
$parser->set_object($handler);
// Set the handlers
$parser->set_element_handler('openHandler','closeHandler');
$parser->set_data_handler('dataHandler');
$parser->set_escape_handler('escapeHandler');
// Parse the document
$parser->parse($doc);
$memory[] = memory_get_usage();
unset($parser);
$memory[] = memory_get_usage();
// echo ( $handler->getXHTML() );
echo "<pre>";
print_r($memory);
echo "</pre>";
}
I got these results:
Array
(
[0] => 183432
[1] => 336256
[2] => 340256
[3] => 500264
[4] => 500264
)
Array
(
[0] => 500264
[1] => 652968
[2] => 654800
[3] => 813984
[4] => 813984
)
Array
(
[0] => 813984
[1] => 966688
[2] => 968520
[3] => 1127712
[4] => 1127712
)
etc.
So it is clear that PHP allocate about 150kB (what is size of file) of memory after reading HTML file to $doc variable and then another 150kB after parsing file for new XHTML file (which is stored within object's variable $xhtml). It's logical but how to free this memory after I don't need neither $doc variable nor parser object? Unsetting the variable or object doesn't help. Thanks,
Lubos