What is http://google.com/unclesam ? Edit: must have just mistyped it, works fine.
Here is the robots.txt of Google
21–30 of 34 posts
Re: Here is the robots.txt of Google
#22Re: Here is the robots.txt of Google
#23http://yahoo.com/robots.txt Sorry, the page you requested was not found.
Re: Here is the robots.txt of Google
#24Earlier quoted context omitted.
Though you could end up with a load of URLs, titles, and a summary text, crawling google won't really get you to the important data in building a search engine. No follow-on links, only a subset of the text, no referrals, etc. Love the 'paradox' of the infinite spider loop, hadn't considered that before.
Crawling groups would be interesting. What other archives of usenet are there?
Re: Here is the robots.txt of Google
#25What is http://google.com/unclesam ? Edit: must have just mistyped it, works fine.
it searches the .gov domain, because 'uncle sam' employees can't remember to put site:.gov in front of their searches ? Think of it as a shortcut.
Re: Here is the robots.txt of Google
#26Earlier quoted context omitted.
Yes, but... For the most part "Google" is a condensed version of the web. In theory, you could "spider" a single site (google.com) and build your own search database without having to go out and crawl the web at large. It's an odd paradox, but I think one search site crawling another search site is not a good idea. And there is probably an infinite spider loop hiding in that process somewhere.
Though you could end up with a load of URLs, titles, and a summary text, crawling google won't really get you to the important data in building a search engine. No follow-on links, only a subset of the text, no referrals, etc. Love the 'paradox' of the infinite spider loop, hadn't considered that before.
Re: Here is the robots.txt of Google
#27Does Yahoo! not have a robots.txt file? http://yahoo.com/robots.txt Sorry, the page you requested was not found.
Re: Here is the robots.txt of Google
#28I always figured that the company that spiders everybody elses content should have a more relaxed policy towards being spidered itself. After all, google is datamining the web on an ongoing basis, in return it should willingly consent to being mined in return.
Yes, but... For the most part "Google" is a condensed version of the web. In theory, you could "spider" a single site (google.com) and build your own search database without having to go out and crawl the web at large. It's an odd paradox, but I think one search site crawling another search site is not a good idea. And there is probably an infinite spider loop hiding in that process somewhere.
Re: Here is the robots.txt of Google
#29Does Yahoo! not have a robots.txt file? http://yahoo.com/robots.txt Sorry, the page you requested was not found.
If that page doesn't exist, then according to specification, they don't and you can crawl any page.
Yes:
http://search.yahoo.com/robots.txt
http://groups.yahoo.com/robots.txt
http://realestate.yahoo.com/robots.txt
No: http://maps.yahoo.com/robots.txt
http://omg.yahoo.com/robots.txt