Earlier quoted context omitted.
The Internet Archive choosing to honor robots.txt is what's 'banning' the access. Both the request not to be crawled and the decision not to crawl are voluntary, but if the Internet Archive decided it wanted to slurp up the Washington Post tomorrow, there's not much they could do to stop it.
They have explicitly denied permission to have their content slurped. Why do you think it is legal to then go ahead and slurp it?
Robots.txt Disallow: 20 Years of Mistakes To Avoid
21–30 of 63 posts
Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid
#22This article forgot the very worst use of robots.txt: User-agent: ia_archiver Disallow: / Those two lines mean that all content hosted on the entire site will be blocked from the Internet Archive (archive.org) WayBack Machine, and the public will be unable to look at any previous versions of the website's content. It wipes out a public view of the past. Yeah, I'm looking at you, Washington Post: http://www.washington…
Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid
#23Earlier quoted context omitted.
They have explicitly denied permission to have their content slurped. Why do you think it is legal to then go ahead and slurp it?
I'm not making an argument about the legality or ethics of it one way or another - although as far as i'm aware, it's not actually illegal to ignore robots.txt. I was just pointing out that robots.txt doesn't actually do anything but ask nicely.
Politicians keep attempting to write evermore draconian qualifications and punishments into law for what qualifies as a "breach of terms of service". I would expect this to encompass robots.txt at some point if it does not already.
Again, I'm not particularly happy about this trend, but I'll try to keep out of its path of destruction.
Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid
#24My server returns 410 GONE to robots.txt requests. The robots exclusion protocol is a ridiculous anachronism. I don't use it and neither should you.
Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid
#25Earlier quoted context omitted.
The Internet Archive choosing to honor robots.txt is what's 'banning' the access. Both the request not to be crawled and the decision not to crawl are voluntary, but if the Internet Archive decided it wanted to slurp up the Washington Post tomorrow, there's not much they could do to stop it.
Yes, I know, I'm a member of Archive Team, and I use "wget -e robots=off --mirror …" quite a bit, and then I upload those WARC's to the IA. But major content providers like the Washington Post that explicitly choose to block their entire website and its history should be named and shamed. Authors don't get the right to go around removing their novels from public libraries just because they would rather the books be a…
The Internet Archive does wonderful work, but just because somebody doesn't want you folks crawling their content doesn't make them worthy of "naming and shaming"
Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid
#26This article forgot the very worst use of robots.txt: User-agent: ia_archiver Disallow: / Those two lines mean that all content hosted on the entire site will be blocked from the Internet Archive (archive.org) WayBack Machine, and the public will be unable to look at any previous versions of the website's content. It wipes out a public view of the past. Yeah, I'm looking at you, Washington Post: http://www.washington…
No, you do NOT have the right to pound my site with requests and serve data that I decided to pull down.
If it's on their bandwidth and power, why not?
Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid
#27This article forgot the very worst use of robots.txt: User-agent: ia_archiver Disallow: / Those two lines mean that all content hosted on the entire site will be blocked from the Internet Archive (archive.org) WayBack Machine, and the public will be unable to look at any previous versions of the website's content. It wipes out a public view of the past. Yeah, I'm looking at you, Washington Post: http://www.washington…
Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid
#28Earlier quoted context omitted.
The Internet Archive choosing to honor robots.txt is what's 'banning' the access. Both the request not to be crawled and the decision not to crawl are voluntary, but if the Internet Archive decided it wanted to slurp up the Washington Post tomorrow, there's not much they could do to stop it.
Yes, I know, I'm a member of Archive Team, and I use "wget -e robots=off --mirror …" quite a bit, and then I upload those WARC's to the IA. But major content providers like the Washington Post that explicitly choose to block their entire website and its history should be named and shamed. Authors don't get the right to go around removing their novels from public libraries just because they would rather the books be a…
Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid
#29Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid
#30This article forgot the very worst use of robots.txt: User-agent: ia_archiver Disallow: / Those two lines mean that all content hosted on the entire site will be blocked from the Internet Archive (archive.org) WayBack Machine, and the public will be unable to look at any previous versions of the website's content. It wipes out a public view of the past. Yeah, I'm looking at you, Washington Post: http://www.washington…
No, you do NOT have the right to pound my site with requests and serve data that I decided to pull down.