RE: Manuals and Spiders
| From: | Divyank Turakhia | Date: | Tue, 22 Apr 2003 14:52:01 +0000 |
| Subject: | RE: Manuals and Spiders | ||
| References: | 1 | Groups: | php.mirrors |
| Request: | Send a blank email to php-mirrors+get-16892@lists.php.net to get a copy of this message | ||
Hi,
I think the whole idea of forceably putting up a robots.txt is bad. If
you want to do this for your own mirror to save on bandwidth, you could
always put up a robots.txt file yourself (I don't know the php mirror
policy on this, but I don't think they should mind you blocking spiders
from your mirror).
If I, as a user, am searching on the engine of my choice - it could be
google , alltheweb, av - I obviously don't know whether I would actually
get this info that I am searching for in the php manual. A search engine
would give me all the results from multiple sites. If the info is in the
manual, then I wouldn't even know that the option of going to php.net or
its mirrors and doing a search exists.
Added to this, derick rightly pointed out, as of now, its easier to
search on a search engine than use the current search feature on the
website. Also, even if the search does get better on the site, it still
never makes sense to block out spiders.
The only time I may want to block a spider by putting my own robots.txt,
is if I notice some small spider is getting stuck in my site due to bad
code. I would, infact, block out such a spider at my firewall itself.
- Divyank
> -----Original Message-----
> From: Alex Kiesel [mailto:kiesel@schlund.de]
> Sent: Tuesday, April 22, 2003 7:39 PM
> To: php-mirrors@lists.php.net
> Subject: Manuals and Spiders
>
>
> Hi,
>
> looking at the huge percentage of visits that come from
> search robots [1], I think it would be reasonable to create a
> robots.txt that prevents those crawlers to index php.net and
> its mirrors.
>
> php.net itself features a full-text search, many mirrors do
> so, too. PHP does not take any advantage by being searched by
> crawlers (at least in the manuals).
>
> So I'd propose to create a robots.txt that at least prevents
> indexing the manuals.
>
> What do you think?
>
> Cheers,
> -Alex
>
> [1] At this moment, I can grep from my logs:
> > php3:~ # grep 'www.googlebot.com'
> /var/log/httpd/php3.de/access_log | wc -l
> > 989661
> > php3:~ # wc -l /var/log/httpd/php3.de/access_log
> > 9514131 /var/log/httpd/php3.de/access_log
>
> This is roughly 10% of all traffic.
>