Live data from Hacker News

Robots.txt Disallow: 20 Years of Mistakes To Avoid

beussery.com

51–60 of 63 posts

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#51

Earlier quoted context omitted.

I was going to reply pointing out that whether or not to name and shame someone is a subjective decision which you and I do not see eye to eye on, and which generally requires quite a few people to agree with you before it becomes a problem for the shamee, but then I rembered the poor way that IA handles changes in ownership with respect to robots. When IA stops wiping out historical content due to a change of domain…

How is IA supposed to distinguish a new website from a sincere wish to delete old stuff? A change in domain registration data means nothing; I have a domain that I registered for an association in my name, and which I then sold to them (for a symbolic price), but it was only an administrative issue - the site was the same. IA is on iffy territory w.r.t. copyright as it is; if they stop respecting robots.txt, they cou…

Your last sentence is key. As I understand it, there's no real legal precedent for IA which basically copies everything out there on an opt-out basis. I personally am glad they do but one of the ways they get off with it is by treading as lightly as possible, including respecting robots.txt even retroactively.

They're also non-commercial, broad in scope, arguably serve a valuable scholarly function and have other characteristics that have kept them mostly out of legal hot water. But it's unclear to what degree they're legally different from a site that decided to create an archive of all comics, commercial and otherwise, and slap advertising up.

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#52

Earlier quoted context omitted.

"serve data that I decided to pull down." If it's on their bandwidth and power, why not?

I think pekk meant that if he deletes a blog post, the IA is still going to serve it. "Pull down" refers to deleting content, not bandwidth usage.

If you don't want it on the internet, don't post it. Assuming anything can ever be made to disappear from the internet is naive, and if people become aware you are, it'll just get Streisanded and become even more widely posted.

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#53

Earlier quoted context omitted.

I think pekk meant that if he deletes a blog post, the IA is still going to serve it. "Pull down" refers to deleting content, not bandwidth usage.

If you don't want it on the internet, don't post it. Assuming anything can ever be made to disappear from the internet is naive, and if people become aware you are, it'll just get Streisanded and become even more widely posted.

Things disappear from the internet, for good, all the time.

Also, if people become aware someone is naive they automatically are dicks to them? Regardless of wether that information is actually of interest to anyone, just because someone wants to take something down, they should not ever be able to?

Some people act and think like that that, yeah. But to accept this as the baseline of human behaviour is, well, not for me. This entitledness to watch the lives of others from the the dark may have been bred by reality TV or whatever; but it's more a personality flaw and an addiction, a useless misfiring of synapses become culture, than a cornerstone of an information age.

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#54
post #22

Earlier quoted context omitted.

No, you do NOT have the right to pound my site with requests and serve data that I decided to pull down.

Nobody said they do; nobody said the Internet Archive shouldn't respect robots.txt. We do, however, have the right to criticize people who ban IA from their site.

Anybody has the right to criticize anyone; the question rather is, do they have a valid criticism. Your wording being so unclear I don't even know what you think about people banning IA from their site, but assuming you would criticize them, what would that criticism be? And would you also criticize someone for making their site private, or not making a site at all?

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#55

Earlier quoted context omitted.

Nobody said they do; nobody said the Internet Archive shouldn't respect robots.txt. We do, however, have the right to criticize people who ban IA from their site.

Anybody has the right to criticize anyone; the question rather is, do they have a valid criticism. Your wording being so unclear I don't even know what you think about people banning IA from their site, but assuming you would criticize them, what would that criticism be? And would you also criticize someone for making their site private, or not making a site at all?

I think if the site is publicly accessible, it's basic Internet civility to allow IA to archive it, but especially so for a newspaper. It's a question of respect for your users and for journalism.

If they have fears about losing revenue - and although I find them silly - there are other ways of going about it, such as only allowing access to pages some weeks or months after they've been published.

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#56

Earlier quoted context omitted.

I think pekk meant that if he deletes a blog post, the IA is still going to serve it. "Pull down" refers to deleting content, not bandwidth usage.

If you don't want it on the internet, don't post it. Assuming anything can ever be made to disappear from the internet is naive, and if people become aware you are, it'll just get Streisanded and become even more widely posted.

There are two sides to this argument. I argue both. All the fucking time.

If you're in the business of providing public content that's well-known, to the public, then allowing it to be archived makes a lot of sense.

If you're providing user-generated content I'd argue that the case for allowing archival is extended even more so. Sites that violate this, and Quora comes specifically to mind, are violating what many, myself included, consider to be part of the social contract of the Web.

On the other hand, if you're an individual, and you are posting your own content and ramblings, and circumstances change for whatever reason: you've got a job, you've lost a job, you're married, you're divorced, you're getting divorced, your child is at war in a foreign country, a foreign country is at war with yours, or you're just sick of the crap you wrote when you were young and arrogant and now and old and arrogant you wants it gone: I'm pretty willing to grant you that right.

If you've committed some terrible crime against humanity, or just a human, and have been fairly tried and convicted of it, I'd probably not give you the right to remove large bits of that information.

And yes, there are vast fields, deserts, tundras, plains, steppes, ice-fields, and oceans of grey about all of this.

Barbra Streisand got Streisanded because she is Streisand.

Ahmed's Falafel Hut likely wouldn't suffer the same fate. His Q-score is somewhat lower, and there's only so much real estate in the public consciousness.

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#57
post #9
post #6

The article contains some good observations, but I'm struggling to understand this one: "Some sites try to communicate with Google through comments in robots.txt" In the examples given, none appear to be trying to "communicate with Google through comments" - how is including... # What's all this then? # \ # # ----- # | . . | # ----- # \--|-|--/ # | | # |-------| ...a "mistake" to avoid? There's no harm in it at all.

"Some sites try to communicate with Google through comments in robots.txt" I thought that was the whole point of robots.txt

No, the point is to communicate with Google through non-comments in robots.txt.

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#58

What frustrates me is the number of websites that impose additional restrictions on anything they don't recognize, or worse, websites that impose additional restrictions on (or worse yet, just outright ban) anything that isn't Googlebot. And people wonder why alternative search engines have such a hard time taking off.

I can give you a really simple operational reason for that: complexity.

Google is somewhere between 50-90% of most sites' search referrals (source: /dev/ass). Add in a handful of other search engines (Bing, DDG, Yahoo, Ask) and you've pretty much got all of it.

They're maybe 10-20% of your crawl traffic though. And possibly a lot less than that.

There are a TON of bots out there. If you're lucky, they just fill your logs and hammer your bandwidth.

If you're not so lucky, they break your site search, overload your servers, and if you're particularly unlucky, they wake you up with 2:30 am pages for two weeks straight.

At which point the simplest way to solve the technical problem, that is, you getting a full night's sleep, is to ban every last fucking bot but Google. Or maybe a handful of the majors.

Now, of course, you're a data-driven operation and you're relying on Google Analytics to tell you who's sending traffic your way. But if you block a search crawler, it's going to stop sending you traffic, so you won't know it's important.

It's a rather similar set of logic that drives people to set email bans on entire CCTLDs or ASN blocks for foreign countries. And if you're a smallish site, it's probably a decent heuristic. And no, it's not just fucking n00bs who do this. Lauren Weinstein who pretty much personally birthed ARPANET at UCLA was bitching on G+ just a week or so back that the new set of unlimited TLDs ICANN were selling were rapidly going into his mailserver blocklists. Because, of course, the early adoptors of such TLDs tend to be spammers, or at least, the early adopters he's likely to hear from.

https://plus.google.com/114753028665775786510/posts/SsgPNHLG...

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#59
post #25

Earlier quoted context omitted.

It's not really shameworthy to want to regulate access to your own sites, and physical metaphors work about as well here as "you wouldn't steal a car" does for piracy. The Internet Archive does wonderful work, but just because somebody doesn't want you folks crawling their content doesn't make them worthy of "naming and shaming"

Why can a physical library collect and display physical newspapers, but a digital library cannot collect and display digital newspapers?

As I pointed out elsewhere in the thread, making analogies to physical media is just as flawed as the "you wouldn't steal a car" anti-piracy campaigns.

A physical library is either getting their newspapers by asking/paying the newspaper company to deliver them, asking citizens to donate them, or collecting them from already delivered newspapers. If the IA was just piggybacking on user activity (by caching and storing things from a user's browser cache after they visit a page) then I'd have far less of a concern with them. If we're so attached to physical metaphors, this would be equivalent to the librarian running around outside the newspaper's printing room and snatching newspapers from the bundles as the company's employees loaded them onto trucks.

Re: Robots.txt Disallow: 20 Years of Mistakes To Avoid

#60

Earlier quoted context omitted.

I still don't see the relationship between deleting a blog post that you have authored and burning a library full of other people's works down.

The robots.txt will not only disallow your blog post, but if you acquired your domain from someone else, the entire previous site will also be removed. That is not something you should have a right to do, unless you also acquired full copyright to all of the previous site's revisions. So sometimes an IA-friendly domain expires (e.g. accidentally or because its owner died), a squatter buys the domain, and the squatter…

Ah I see, thanks for the explanation
Post reply on HN