Live data from Hacker News

Here is the robots.txt of Google

google.com

31–34 of 34 posts

Re: Here is the robots.txt of Google

#33
post #8

What is http://google.com/unclesam ? Edit: must have just mistyped it, works fine.

this is even weirder.. http://www.google.com/microsoft

It's not weird--it's very useful for looking up Windows support info or IE dev documentation, e.g..

Re: Here is the robots.txt of Google

#34

Earlier quoted context omitted.

Though you could end up with a load of URLs, titles, and a summary text, crawling google won't really get you to the important data in building a search engine. No follow-on links, only a subset of the text, no referrals, etc. Love the 'paradox' of the infinite spider loop, hadn't considered that before.

Most spiders implement a maximum depth in their search though. So you'd really just get an inefficient crawl, not an infinite one.

Of course, using Google as a seed list more than crawling only google. Not sure why I didn't think of that.
Post reply on HN