Re: Parsing HTML documents
| From: | Andy Clarke | Date: | Mon, 04 Dec 2000 13:20:18 +0000 |
| Subject: | Re: Parsing HTML documents | ||
| References: | 1 | Groups: | php.general |
| Request: | Send a blank email to php-general+get-28568@lists.php.net to get a copy of this message | ||
>> I want to write a script that:
>> * Reads a directory on the website
>> * List the names of the HTML files in the directory
>> * Reads the Title tag from each file and generates a hyperlink with it.
>>
>> Is there a simpler way other than manually searching each HTML file for
>the
>> <TITLE> tag?
>Sorry if my answer is a little bit stupid, but don't you think it's easier
>to create pages with the names of the titles you need? Then you could skip
>step 3.
Would it be even easier to simply have no index.html file in the directory?
A listing will be created automatically, and the appearance of this could
be controlled using the Apache mod_autoindex module.
There is a ScanHTMLTitles option, though the documentation warns that this
is "CPU and disk intensive".
ScanHTMLTitles is listed under the IndexOptions directive at:
http://httpd.apache.org/docs/mod/mod_autoindex.html#indexoptions
Andy Clarke
-----------------------------
Andy Clarke
78 West Kensington Court
Edith Villas
London W14 9AB
Phone: 44 (0)20 7602 3382
Mobile: 07947 418177
Email: andy@kinonet.com
-----------------------------