RE: Manuals and Spiders
| From: | Divyank Turakhia | Date: | Tue, 22 Apr 2003 16:44:32 +0000 |
| Subject: | RE: Manuals and Spiders | ||
| References: | 1 | Groups: | php.mirrors |
| Request: | Send a blank email to php-mirrors+get-16904@lists.php.net to get a copy of this message | ||
*Automatic* Mirrors based on the IP address is possible, but it opens up
its own can of worms. We have already had a discussion about automatic
mirroring on this list. In fact, I had stated earlier that we can
provide the ip-country database free for php, all its mirrors and any
other open source intiatives related to the php group. The offer is
still open to use our product - www.ip-to-country.com . We can also put
in some programming resources to help php and its mirrors deploy an
automatic mirror soln.
The main issue that were unresolved at tht time was - wht if the mirror
site is down at tht point of time - then the visitor would keep getting
redirected there. We would need addional code to be written which will
check all the mirrors every 'x' minutes. I don't know of any better soln
to this as yet.
That's as far as the mirroring is concerned, which is a very good idea
and I had had tried to push this earlier too.
As far as blocking spiders is concerned - I still think it's a big NO
NO. Let the user come to any site, the mirroring code can exist at each
site too which will automatically redirect the user to the closest
mirror with the same content. Blocking search engines from crawling your
website is, as such, a bad idea. The engine is supposed to figure out,
how it should make sure that the main php website is listed before the
other mirrors. Google already does this using its PR formulas. I have no
idea how other engines do this, but this is not out concern any which
ways. We have content on our site, if their engine wants to link to it,
they can do so however they want. I see no reason for blocking them.
- Divyank
> -----Original Message-----
> From: shimi [mailto:shimi@shimi.net]
> Sent: Tuesday, April 22, 2003 8:28 PM
> To: Divyank Turakhia
> Cc: 'Alex Kiesel'; php-mirrors@lists.php.net
> Subject: RE: Manuals and Spiders
>
>
>
> I personally think that all mirrors should block all spiders,
> except for
> www.
>
> Why? Because I am constantly getting linked from Google from
> people from
> Taiwan, etc, to my mirror (which is on the other side of earth...)
>
> This is not smart. Only the www links should appear in search
> engines (in
> my opinion), and the main www should be mirrors aware, and
> direct people
> *automatically* to mirrors, thus, always giving users:
>
> 1) fast php net
>
> 2) take off most of the load off www.php.net's server.
>
> my 2 cents ;)
>
> On Tue, 22 Apr 2003, Divyank Turakhia wrote:
>
> > Hi,
> >
> > I think the whole idea of forceably putting up a robots.txt
> is bad. If
> > you want to do this for your own mirror to save on bandwidth, you
> > could always put up a robots.txt file yourself (I don't
> know the php
> > mirror policy on this, but I don't think they should mind
> you blocking
> > spiders from your mirror).
> >
> > If I, as a user, am searching on the engine of my choice -
> it could be
> > google , alltheweb, av - I obviously don't know whether I would
> > actually get this info that I am searching for in the php manual. A
> > search engine would give me all the results from multiple sites. If
> > the info is in the manual, then I wouldn't even know that
> the option
> > of going to php.net or its mirrors and doing a search exists.
> >
> > Added to this, derick rightly pointed out, as of now, its easier to
> > search on a search engine than use the current search
> feature on the
> > website. Also, even if the search does get better on the site, it
> > still never makes sense to block out spiders.
> >
> > The only time I may want to block a spider by putting my own
> > robots.txt, is if I notice some small spider is getting stuck in my
> > site due to bad code. I would, infact, block out such a
> spider at my
> > firewall itself.
> >
> > - Divyank
> >
> > > -----Original Message-----
> > > From: Alex Kiesel [mailto:kiesel@schlund.de]
> > > Sent: Tuesday, April 22, 2003 7:39 PM
> > > To: php-mirrors@lists.php.net
> > > Subject: Manuals and Spiders
> > >
> > >
> > > Hi,
> > >
> > > looking at the huge percentage of visits that come from
> > > search robots [1], I think it would be reasonable to create a
> > > robots.txt that prevents those crawlers to index php.net and
> > > its mirrors.
> > >
> > > php.net itself features a full-text search, many mirrors do
> > > so, too. PHP does not take any advantage by being searched by
> > > crawlers (at least in the manuals).
> > >
> > > So I'd propose to create a robots.txt that at least prevents
> > > indexing the manuals.
> > >
> > > What do you think?
> > >
> > > Cheers,
> > > -Alex
> > >
> > > [1] At this moment, I can grep from my logs:
> > > > php3:~ # grep 'www.googlebot.com'
> > > /var/log/httpd/php3.de/access_log | wc -l
> > > > 989661
> > > > php3:~ # wc -l /var/log/httpd/php3.de/access_log
> > > > 9514131 /var/log/httpd/php3.de/access_log
> > >
> > > This is roughly 10% of all traffic.
> > >
> >
>
> --
>
> Best regards,
> Shimi
>
>
> ----
>
> "Outlook is a massive flaming horrid blatant security
> violation, which
> also happens to be a mail reader."
>
> "Sure UNIX is user friendly; it's just picky about who its
> friends are."
>
>