RE: Manuals and Spiders

From: Date: Tue, 22 Apr 2003 14:52:01 +0000
Subject: RE: Manuals and Spiders
References: 1  Groups: php.mirrors 
Request: Send a blank email to php-mirrors+get-16892@lists.php.net to get a copy of this message
Hi, I think the whole idea of forceably putting up a robots.txt is bad. If you want to do this for your own mirror to save on bandwidth, you could always put up a robots.txt file yourself (I don't know the php mirror policy on this, but I don't think they should mind you blocking spiders from your mirror). If I, as a user, am searching on the engine of my choice - it could be google , alltheweb, av - I obviously don't know whether I would actually get this info that I am searching for in the php manual. A search engine would give me all the results from multiple sites. If the info is in the manual, then I wouldn't even know that the option of going to php.net or its mirrors and doing a search exists. Added to this, derick rightly pointed out, as of now, its easier to search on a search engine than use the current search feature on the website. Also, even if the search does get better on the site, it still never makes sense to block out spiders. The only time I may want to block a spider by putting my own robots.txt, is if I notice some small spider is getting stuck in my site due to bad code. I would, infact, block out such a spider at my firewall itself. - Divyank > -----Original Message----- > From: Alex Kiesel [mailto:kiesel@schlund.de] > Sent: Tuesday, April 22, 2003 7:39 PM > To: php-mirrors@lists.php.net > Subject: Manuals and Spiders > > > Hi, > > looking at the huge percentage of visits that come from > search robots [1], I think it would be reasonable to create a > robots.txt that prevents those crawlers to index php.net and > its mirrors. > > php.net itself features a full-text search, many mirrors do > so, too. PHP does not take any advantage by being searched by > crawlers (at least in the manuals). > > So I'd propose to create a robots.txt that at least prevents > indexing the manuals. > > What do you think? > > Cheers, > -Alex > > [1] At this moment, I can grep from my logs: > > php3:~ # grep 'www.googlebot.com' > /var/log/httpd/php3.de/access_log | wc -l > > 989661 > > php3:~ # wc -l /var/log/httpd/php3.de/access_log > > 9514131 /var/log/httpd/php3.de/access_log > > This is roughly 10% of all traffic. >

« previous php.mirrors (#16892) next »