Live data from Hacker News

Here is the robots.txt of Google

google.com

1–10 of 34 posts

Re: Here is the robots.txt of Google

#2
I always figured that the company that spiders everybody elses content should have a more relaxed policy towards being spidered itself.

After all, google is datamining the web on an ongoing basis, in return it should willingly consent to being mined in return.

Re: Here is the robots.txt of Google

#3
post #2

I always figured that the company that spiders everybody elses content should have a more relaxed policy towards being spidered itself. After all, google is datamining the web on an ongoing basis, in return it should willingly consent to being mined in return.

Yes, but...

For the most part "Google" is a condensed version of the web. In theory, you could "spider" a single site (google.com) and build your own search database without having to go out and crawl the web at large.

It's an odd paradox, but I think one search site crawling another search site is not a good idea. And there is probably an infinite spider loop hiding in that process somewhere.

Re: Here is the robots.txt of Google

#4
post #3
post #2

I always figured that the company that spiders everybody elses content should have a more relaxed policy towards being spidered itself. After all, google is datamining the web on an ongoing basis, in return it should willingly consent to being mined in return.

Yes, but... For the most part "Google" is a condensed version of the web. In theory, you could "spider" a single site (google.com) and build your own search database without having to go out and crawl the web at large. It's an odd paradox, but I think one search site crawling another search site is not a good idea. And there is probably an infinite spider loop hiding in that process somewhere.

Not like anyone writing a spider has to obey the robots.txt file.

Re: Here is the robots.txt of Google

#5
post #2

I always figured that the company that spiders everybody elses content should have a more relaxed policy towards being spidered itself. After all, google is datamining the web on an ongoing basis, in return it should willingly consent to being mined in return.

I think the goal is to prevent other robots from accidentally crawling the Google corpus. At the end of the day, robots.txt is a stab at comity. There are many well-known services which use crawlers that fail to obey it (or identify themselves). The file merely announces, "This is all recycled content." -- notice that none of the Google corporate pages are disallowed.

Aside: let's say you've whipped up a spiffy new ranking algorithm, and you just need an index to launch your search engine. What's faster: crawling the web, or crawling Google? I don't think such an entrepreneur would pass on a big speed up just because of a text-file.

Re: Here is the robots.txt of Google

#6
post #3
post #2

I always figured that the company that spiders everybody elses content should have a more relaxed policy towards being spidered itself. After all, google is datamining the web on an ongoing basis, in return it should willingly consent to being mined in return.

Yes, but... For the most part "Google" is a condensed version of the web. In theory, you could "spider" a single site (google.com) and build your own search database without having to go out and crawl the web at large. It's an odd paradox, but I think one search site crawling another search site is not a good idea. And there is probably an infinite spider loop hiding in that process somewhere.

And there is probably an infinite spider loop hiding in that process somewhere.

And then, this strange loop attains self awareness... (GEB reference)

Re: Here is the robots.txt of Google

#7
post #4
post #3

Earlier quoted context omitted.

Yes, but... For the most part "Google" is a condensed version of the web. In theory, you could "spider" a single site (google.com) and build your own search database without having to go out and crawl the web at large. It's an odd paradox, but I think one search site crawling another search site is not a good idea. And there is probably an infinite spider loop hiding in that process somewhere.

Not like anyone writing a spider has to obey the robots.txt file.

I smell a business opportunity here! I could develop a Digital Robots.txt Management platform that will be built into all server operating systems and would enforce that no spiders are allowed to disregards robots.txt. I'm a genius! Oh, wait...

Re: Here is the robots.txt of Google

#10
post #4
post #3

Earlier quoted context omitted.

Yes, but... For the most part "Google" is a condensed version of the web. In theory, you could "spider" a single site (google.com) and build your own search database without having to go out and crawl the web at large. It's an odd paradox, but I think one search site crawling another search site is not a good idea. And there is probably an infinite spider loop hiding in that process somewhere.

Not like anyone writing a spider has to obey the robots.txt file.

Sure, and their IP will be blocked by Google some time around the 10th request.
Post reply on HN