Re: Robots.txt file restricting msnbot at http://news.php.net/robots.txt
| From: | Hannes Magnusson | Date: | Wed, 22 Oct 2008 06:54:04 +0000 |
| Subject: | Re: Robots.txt file restricting msnbot at http://news.php.net/robots.txt | ||
| References: | 1 2 3 | Groups: | php.webmaster |
| Request: | Send a blank email to php-webmaster+get-2843@lists.php.net to get a copy of this message | ||
On Wed, Oct 22, 2008 at 01:41, Amy Wilcox (Murphy & Associates)
<v-amwilc@microsoft.com> wrote:
> Hi Hannes,
>
>
>
> I apologize for the delayed response. Your robots.txt at
> http://news.php.net/robots.txt is preventing Live Search
> (msnbot) from
> indexing your site but allowing other spiders (Slurp).
Errr... Why do you want to index our news frontend anyway? You surely
have other means of indexing our mailinglist archives, don't you?
The reason why all spiders where blocked was do to serious spider
traffic on the server, as you can imagine since the archives over the
past 10 years or so have gotten huge. I honestly do not see why you
want to index our mailinglist frontend.. it is publically archived by
dozens of mailinglist archives out there, and you are totally free to
archive it too yourself - but indexing our (already extremely slow)
http frontend to the mailinglists seems totally worthless to me.
Why Slurp is allowed has to do with our relationship with Y! and their
search engineers, but that was really more of an experiment then a
serious "please crawl our http mailinglist archives" invitation.
The same thing applies to the http frontend to our CVS repository,
http://cvs.php.net, which we don't even allow Slurp to index
as there
is absolutely no point in it.
-Hannes