Re: Metasearches
| From: | Felipe G. Coury | Date: | Tue, 10 Oct 2000 19:00:57 +0000 |
| Subject: | Re: Metasearches | ||
| References: | 1 | Groups: | php.general |
| Request: | Send a blank email to php-general+get-19395@lists.php.net to get a copy of this message | ||
Kenny,
> A good HTML parser may take monthes to create. Since the page
> you get might not have good HTML style, which mean, a lot of tags
> are not closed properly, your parser needs tolerate all that, and still
> generate
> useful indexes. A good one should take about 1 year to develop and
> testing, and might only work with the pages that doesn't have
> javascripts in the links.
Thanks for making that clear for me.
> Or, you can try to use the XML parser in the PHP to parse the HTML
> pages. Again, since XML parser will halt if the input page is not
> properly written, your parsing result might not be complete if the
> page you have contains unclosed tags, miss placement of tags,...etc.
> This parser doesn't work with the javascript too. But if you know
> clearly about the structure of page you want to parse, you can use
> ereg functions to extract part of the page that contains the data
> you need, and then use XML parser to process this part of page
> ( if it has good style), you might get what you wanted very quick.
Do you have any resources on how can I learn Regular Expresions in a quick
time (at lease faster time)?
> By the way, you should also take a look at the robot standard
> if you are trying to run a robot on someone's search engine or
> webpages.
I'll be reviewing this.
Regards,
Felipe G. Coury
Gerente de Desenvolvimento
Creation Internet Business
http://www.creation.com.br
--------------------------------------------
Visite o GuiaBrasil.Net:
http://www.guiabrasil.net
--------------------------------------------