Live data from Hacker News

Robots.txt Disallow: 20 Years of Mistakes To Avoid

beussery.com

21–30 of 63 posts

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#21
post #18
post #11

Earlier quoted context omitted.

The Internet Archive choosing to honor robots.txt is what's 'banning' the access. Both the request not to be crawled and the decision not to crawl are voluntary, but if the Internet Archive decided it wanted to slurp up the Washington Post tomorrow, there's not much they could do to stop it.

They have explicitly denied permission to have their content slurped. Why do you think it is legal to then go ahead and slurp it?

I'm not making an argument about the legality or ethics of it one way or another - although as far as i'm aware, it's not actually illegal to ignore robots.txt. I was just pointing out that robots.txt doesn't actually do anything but ask nicely.

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#22

This article forgot the very worst use of robots.txt: User-agent: ia_archiver Disallow: / Those two lines mean that all content hosted on the entire site will be blocked from the Internet Archive (archive.org) WayBack Machine, and the public will be unable to look at any previous versions of the website's content. It wipes out a public view of the past. Yeah, I'm looking at you, Washington Post: http://www.washington…

No, you do NOT have the right to pound my site with requests and serve data that I decided to pull down.

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#23
post #21
post #18

Earlier quoted context omitted.

They have explicitly denied permission to have their content slurped. Why do you think it is legal to then go ahead and slurp it?

I'm not making an argument about the legality or ethics of it one way or another - although as far as i'm aware, it's not actually illegal to ignore robots.txt. I was just pointing out that robots.txt doesn't actually do anything but ask nicely.

When people can go to jail for hitting a publicly available URL, I'd question the "legality" of such activity. (I'm not making a moral argument, but rather question what lawyers and law enforcement may choose to make of a situation.)

Politicians keep attempting to write evermore draconian qualifications and punishments into law for what qualifies as a "breach of terms of service". I would expect this to encompass robots.txt at some point if it does not already.

Again, I'm not particularly happy about this trend, but I'll try to keep out of its path of destruction.

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#25
post #11

Earlier quoted context omitted.

The Internet Archive choosing to honor robots.txt is what's 'banning' the access. Both the request not to be crawled and the decision not to crawl are voluntary, but if the Internet Archive decided it wanted to slurp up the Washington Post tomorrow, there's not much they could do to stop it.

Yes, I know, I'm a member of Archive Team, and I use "wget -e robots=off --mirror …" quite a bit, and then I upload those WARC's to the IA. But major content providers like the Washington Post that explicitly choose to block their entire website and its history should be named and shamed. Authors don't get the right to go around removing their novels from public libraries just because they would rather the books be a…

It's not really shameworthy to want to regulate access to your own sites, and physical metaphors work about as well here as "you wouldn't steal a car" does for piracy.

The Internet Archive does wonderful work, but just because somebody doesn't want you folks crawling their content doesn't make them worthy of "naming and shaming"

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#26
post #22

This article forgot the very worst use of robots.txt: User-agent: ia_archiver Disallow: / Those two lines mean that all content hosted on the entire site will be blocked from the Internet Archive (archive.org) WayBack Machine, and the public will be unable to look at any previous versions of the website's content. It wipes out a public view of the past. Yeah, I'm looking at you, Washington Post: http://www.washington…

No, you do NOT have the right to pound my site with requests and serve data that I decided to pull down.

"serve data that I decided to pull down."

If it's on their bandwidth and power, why not?

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#27

This article forgot the very worst use of robots.txt: User-agent: ia_archiver Disallow: / Those two lines mean that all content hosted on the entire site will be blocked from the Internet Archive (archive.org) WayBack Machine, and the public will be unable to look at any previous versions of the website's content. It wipes out a public view of the past. Yeah, I'm looking at you, Washington Post: http://www.washington…

The thing that really frustrates me about the Internet Archive's treatment of robots.txt: if a domain expires and the domain provider changes the robots.txt to something restrictive, the Wayback Machine will completely clear the history of the site. Even though it's very clearly not the same agent at play-- this is not the creator of the site's content. I've seen it happen, and it breaks my heart every time.

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#28
post #11

Earlier quoted context omitted.

The Internet Archive choosing to honor robots.txt is what's 'banning' the access. Both the request not to be crawled and the decision not to crawl are voluntary, but if the Internet Archive decided it wanted to slurp up the Washington Post tomorrow, there's not much they could do to stop it.

Yes, I know, I'm a member of Archive Team, and I use "wget -e robots=off --mirror …" quite a bit, and then I upload those WARC's to the IA. But major content providers like the Washington Post that explicitly choose to block their entire website and its history should be named and shamed. Authors don't get the right to go around removing their novels from public libraries just because they would rather the books be a…

Unfair to name & shame a private entity that doesn't want it's content to be archived.

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#30
post #22

This article forgot the very worst use of robots.txt: User-agent: ia_archiver Disallow: / Those two lines mean that all content hosted on the entire site will be blocked from the Internet Archive (archive.org) WayBack Machine, and the public will be unable to look at any previous versions of the website's content. It wipes out a public view of the past. Yeah, I'm looking at you, Washington Post: http://www.washington…

No, you do NOT have the right to pound my site with requests and serve data that I decided to pull down.

Do you also want to steal into libraries in the night and set fire to their microfiche collection?
Post reply on HN