Live data from Hacker News

Robots.txt Disallow: 20 Years of Mistakes To Avoid

beussery.com

41–50 of 63 posts

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#41
post #25

Earlier quoted context omitted.

Yes, I know, I'm a member of Archive Team, and I use "wget -e robots=off --mirror …" quite a bit, and then I upload those WARC's to the IA. But major content providers like the Washington Post that explicitly choose to block their entire website and its history should be named and shamed. Authors don't get the right to go around removing their novels from public libraries just because they would rather the books be a…

It's not really shameworthy to want to regulate access to your own sites, and physical metaphors work about as well here as "you wouldn't steal a car" does for piracy. The Internet Archive does wonderful work, but just because somebody doesn't want you folks crawling their content doesn't make them worthy of "naming and shaming"

I was going to reply pointing out that whether or not to name and shame someone is a subjective decision which you and I do not see eye to eye on, and which generally requires quite a few people to agree with you before it becomes a problem for the shamee, but then I rembered the poor way that IA handles changes in ownership with respect to robots.

When IA stops wiping out historical content due to a change of domain ownership in the now then I will have more support (and USE) for them.

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#42
post #22

This article forgot the very worst use of robots.txt: User-agent: ia_archiver Disallow: / Those two lines mean that all content hosted on the entire site will be blocked from the Internet Archive (archive.org) WayBack Machine, and the public will be unable to look at any previous versions of the website's content. It wipes out a public view of the past. Yeah, I'm looking at you, Washington Post: http://www.washington…

No, you do NOT have the right to pound my site with requests and serve data that I decided to pull down.

I may not have a right to "pound" your site, but I certainly have a right to keep whatever I find on your public webserver, regardless of whether you decide to pull it down later.

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#43
post #35

Earlier quoted context omitted.

I'm assuming this was in response to the "and serve data that I decided to pull down." part.

I still don't see the relationship between deleting a blog post that you have authored and burning a library full of other people's works down.

The robots.txt will not only disallow your blog post, but if you acquired your domain from someone else, the entire previous site will also be removed. That is not something you should have a right to do, unless you also acquired full copyright to all of the previous site's revisions.

So sometimes an IA-friendly domain expires (e.g. accidentally or because its owner died), a squatter buys the domain, and the squatter points the domain at a junk-site landing server with a deny all robots.txt. The result is truly disastrous: IA removes access to the historic, IA-friendly site. Site acquirers who do this deliberately are pure evil.

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#44
post #25

Earlier quoted context omitted.

Yes, I know, I'm a member of Archive Team, and I use "wget -e robots=off --mirror …" quite a bit, and then I upload those WARC's to the IA. But major content providers like the Washington Post that explicitly choose to block their entire website and its history should be named and shamed. Authors don't get the right to go around removing their novels from public libraries just because they would rather the books be a…

It's not really shameworthy to want to regulate access to your own sites, and physical metaphors work about as well here as "you wouldn't steal a car" does for piracy. The Internet Archive does wonderful work, but just because somebody doesn't want you folks crawling their content doesn't make them worthy of "naming and shaming"

Why can a physical library collect and display physical newspapers, but a digital library cannot collect and display digital newspapers?

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#45
post #22

This article forgot the very worst use of robots.txt: User-agent: ia_archiver Disallow: / Those two lines mean that all content hosted on the entire site will be blocked from the Internet Archive (archive.org) WayBack Machine, and the public will be unable to look at any previous versions of the website's content. It wipes out a public view of the past. Yeah, I'm looking at you, Washington Post: http://www.washington…

No, you do NOT have the right to pound my site with requests and serve data that I decided to pull down.

Nobody said they do; nobody said the Internet Archive shouldn't respect robots.txt.

We do, however, have the right to criticize people who ban IA from their site.

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#46
post #25

Earlier quoted context omitted.

It's not really shameworthy to want to regulate access to your own sites, and physical metaphors work about as well here as "you wouldn't steal a car" does for piracy. The Internet Archive does wonderful work, but just because somebody doesn't want you folks crawling their content doesn't make them worthy of "naming and shaming"

I was going to reply pointing out that whether or not to name and shame someone is a subjective decision which you and I do not see eye to eye on, and which generally requires quite a few people to agree with you before it becomes a problem for the shamee, but then I rembered the poor way that IA handles changes in ownership with respect to robots. When IA stops wiping out historical content due to a change of domain…

How is IA supposed to distinguish a new website from a sincere wish to delete old stuff? A change in domain registration data means nothing; I have a domain that I registered for an association in my name, and which I then sold to them (for a symbolic price), but it was only an administrative issue - the site was the same.

IA is on iffy territory w.r.t. copyright as it is; if they stop respecting robots.txt, they could get into a world of hurt.

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#47

This article forgot the very worst use of robots.txt: User-agent: ia_archiver Disallow: / Those two lines mean that all content hosted on the entire site will be blocked from the Internet Archive (archive.org) WayBack Machine, and the public will be unable to look at any previous versions of the website's content. It wipes out a public view of the past. Yeah, I'm looking at you, Washington Post: http://www.washington…

The thing that really frustrates me about the Internet Archive's treatment of robots.txt: if a domain expires and the domain provider changes the robots.txt to something restrictive, the Wayback Machine will completely clear the history of the site. Even though it's very clearly not the same agent at play-- this is not the creator of the site's content. I've seen it happen, and it breaks my heart every time.

Why wouldn't it consider the archived state of robots.txt?

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#49

This article forgot the very worst use of robots.txt: User-agent: ia_archiver Disallow: / Those two lines mean that all content hosted on the entire site will be blocked from the Internet Archive (archive.org) WayBack Machine, and the public will be unable to look at any previous versions of the website's content. It wipes out a public view of the past. Yeah, I'm looking at you, Washington Post: http://www.washington…

The thing that really frustrates me about the Internet Archive's treatment of robots.txt: if a domain expires and the domain provider changes the robots.txt to something restrictive, the Wayback Machine will completely clear the history of the site. Even though it's very clearly not the same agent at play-- this is not the creator of the site's content. I've seen it happen, and it breaks my heart every time.

One of the reasons I like archive.today. Obviously, they lack the depth of history, but they don't censor so easily.

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#50

Earlier quoted context omitted.

"serve data that I decided to pull down." If it's on their bandwidth and power, why not?

I think pekk meant that if he deletes a blog post, the IA is still going to serve it. "Pull down" refers to deleting content, not bandwidth usage.

If you don't want it on the internet, don't post it. Assuming anything can ever be made to disappear from the internet is naive, and if people become aware you are, it'll just get Streisanded and become even more widely posted.
Post reply on HN