Live data from Hacker News

Google's robots.txt

google.com

91–94 of 94 posts

Re: Google's robots.txt

#91
post #31

http://www.google.com/baraza/en What a weird little product. It's like Yahoo Answers, but somehow with even less sorting or categorization.

Amazingly, the questions seem to be even dumber than the ones found on YA:

- why is computer an idiot machine

- i casted a love spell on my ex,should i tell her now that she is back

- What is the colour of the black box which is using in planes ?

Re: Google's robots.txt

#92
post #15

Earlier quoted context omitted.

I've found this GitHub project to be an invaluable resource for blocking bad bots: https://github.com/bluedragonz/bad-bot-blocker

>Unless your website is written in Russian or Chinese, you probably don't get any traffic from them. They mostly just waste bandwidth and consume resources. THIS is evil. You could use this argument for banning any new search engine.

The problem is that both Yandex and Baidu are rather poorly behaved - they hit your website way too fast, downloading large bandwidth files in quick succession. That's actually what led me to the bad-bot-blocker project in the first place. Baidu has also been accused of not respecting robots.txt though I have not personally observed that.

This is the reason they're blocked, not because they're new or non-English.

Re: Google's robots.txt

#93
post #60

Earlier quoted context omitted.

Why you consider Yandex evil?

I don't know if this is universal, but Yandex was quite bad in the past at stumbling around generated links (E.G. calendars that go to infinity) and it wasn't at all unusual to see 400+ links crawled in a day or so. Baidu was the same, but I think they're more well behaved these days.

Lucky you with only 400+ per day. We had thousands per hour. Still the server survived but it wasn't fast...
Post reply on HN