Re: Robots.txt file restricting msnbot at http://news.php.net/robots.txt

From: Date: Wed, 22 Oct 2008 06:54:04 +0000
Subject: Re: Robots.txt file restricting msnbot at http://news.php.net/robots.txt
References: 1 2 3  Groups: php.webmaster 
Request: Send a blank email to php-webmaster+get-2843@lists.php.net to get a copy of this message
On Wed, Oct 22, 2008 at 01:41, Amy Wilcox (Murphy & Associates) <v-amwilc@microsoft.com> wrote: > Hi Hannes, > > > > I apologize for the delayed response. Your robots.txt at > http://news.php.net/robots.txt is preventing Live Search > (msnbot) from > indexing your site but allowing other spiders (Slurp). Errr... Why do you want to index our news frontend anyway? You surely have other means of indexing our mailinglist archives, don't you? The reason why all spiders where blocked was do to serious spider traffic on the server, as you can imagine since the archives over the past 10 years or so have gotten huge. I honestly do not see why you want to index our mailinglist frontend.. it is publically archived by dozens of mailinglist archives out there, and you are totally free to archive it too yourself - but indexing our (already extremely slow) http frontend to the mailinglists seems totally worthless to me. Why Slurp is allowed has to do with our relationship with Y! and their search engineers, but that was really more of an experiment then a serious "please crawl our http mailinglist archives" invitation. The same thing applies to the http frontend to our CVS repository, http://cvs.php.net, which we don't even allow Slurp to index as there is absolutely no point in it. -Hannes

« previous php.webmaster (#2843) next »