Live data from Hacker News

Tell HN: Internet Archive is facing a Big 4 Publishers lawsuit

news.ycombinator.com

211–220 of 226 posts

Re: Tell HN: Internet Archive is facing a Big 4 Publishers lawsuit

#211
post #210
post #209

Earlier quoted context omitted.

The internet has given everyone the tools to publish a personal newsletter/blog/profile-page, just as only the few could do so earlier. It not only allows essentially-free publishing to the whole public, but also extreme narrowcasting with any level of access-control one desires. You've not made a case that personal writings should be treated any differently, on either copyright or privacy grounds, nor that the law d…

> How heavenly it might be if only every paid employee (and potentially, compensated advocate) of big tech, big copyright, big nation-state, big regulation, and big ideologies – as they pile-on the votes & comments here & elsewhere – were similarly open about their affiliations! Luckily I'm not paid to talk about privacy. While there's probably some people who monetize privacy they're usually looked at negatively. Ap…

Your other link (https://jonathanwthomas.net/how-to-get-your-website-out-of-t...) explicitly and accurately reported, "You don’t need to file a DMCA notice or anything (but you can if you want to go nuclear)." Your insistence on sloppily referring to the process as such, and further your unwarranted fear it might "change on a dime", are odd slurs to insistently deploy.

> Doesn't really matter, disclose your affiliations - especially for the kind of wild statements you make (eg: comparing thread commenters to "big tech, big copyright, big nation-state, big regulation, and big ideologies".)

My current & prior affiliations are well-disclosed – better than yours, it seems.

And I wasn't "comparing" commenters to those things, I was pointing out: many pseudonymous voters, & commenters, here are often in the employ of, & de facto advocates for, the very topics of contentious discussion – such as Google, the US federal government, the entertainment industry, activist organizations, etc – with no disclosure.

So to harp on my open book of work history is again odd.

> I agree, commons is the wrong term, "fair use" is what IA legally rides on.

Indeed, it was unfair of you to characterize my arguments, or separately any rationales used by the Internet Archive, as involving a simplistic 'commons' assertion.

> I've also proposed ways to [make their services less harmful] that you have not responded to.

I've not noticed these proposals, just vague platitudes about "trying to establish consensus in good faith first" – what does that mean? – or an example of one person (Manley) who consciously chose to publish a life-archive along with their suicide note (!). How exactly does that outlier case convert to actionable policies for a web library, that legitimately seeks to be as comprehensive about publicly-published web materials as traditional printed-material libraries?

> Regulation will solve anyone who wants to host or do business on US soil.

No, it can't 'solve', and barely even helps, because the most serious threats are from government who themselves use regulations to force the violation of privacy – as with KYC rules, or demands for intercept capabilities – and from aggressive sub-state actors whose activites are largely invisible to regulators.

You're still alleging non-specific "harm" from the Archive without examples or magnitudes.

If the Archive helps makes people aware that what they've published persists in the public record unless they take conscious other steps, and lets them both correct the storage that surprised them, and leads them towards the truly-reliable practices for avoiding unwanted disclosure/persistence, it's done a better job than those promising safety from superficial 'regulation' that doesn't actually limit most threats.

Archiving publically-offered web pages isn't a "privacy-eroding tool" no matter how many times you repeat that allegation - it's a tool for cultural memory, honest history, and teaching people the ground realities of privacy (or lack thereof) in these new media. Directing ire at the Archive's well-behaved collection activities is getting angry at the smoke alarm, not the fire.

Re: Tell HN: Internet Archive is facing a Big 4 Publishers lawsuit

#212

Earlier quoted context omitted.

Thank you for finding this! Still three times less than a Kindle Unlimited subscription :)

I'd guess the Kindle is more convenient, no waiting for a title, no late fees, and has a vastly larger selection (on average). Plus one can put their own content on Kindle, thereby taking lots of titles on trips or elsewhere.

Yup. Convenience is hard to beat.

Re: Tell HN: Internet Archive is facing a Big 4 Publishers lawsuit

#213
post #15

Earlier quoted context omitted.

The mission of preserving human culture is far more important than respecting rent-seeking copyright holders. At the end of the day, The Internet Archive has good intentions and is morally in the right. The time has come to consider changing the laws to allow for truly fair use, especially for physical items scanned to digital (e.g. books), old video games, and more. It's about selecting for the common good over the…

> The mission of preserving human culture is far more important than respecting rent-seeking copyright holders That genuinely may be so, but "this law i broke shouldn't exist" is not an advisable legal defense. Also, what you call "rent seeking" others would call "return on investment". I do think there is a grey area here, "fair use" being one example, but i think summarily discounting distributors and publishers wh…

"this law i broke shouldn't exist"

Wrong. Ultra-vires laws can be challenged after the fact that they were broken. If the law in itself was invalid, that is a valid legal defence.

Re: Tell HN: Internet Archive is facing a Big 4 Publishers lawsuit

#214
post #208
post #200

Earlier quoted context omitted.

In truth, regulation can only reliably protect your privacy from well-behaved actors whose actions/violations are observable. If you've taken no self-help measures to limit access, then bad actors, unobservable to you and regulators, will still be doing whatever they would like to do and can get away with. But you may be lulled into a false sense of security by the false promise of a 'solution' via regulations.

As I've now learned, you used to work for the Internet Archive. You should probably start your statements with that. > If you've taken no self-help measures to limit access... robots.txt was a nice self-help measure. > ... then bad actors, unobservable to you and regulators, will still be doing whatever they would like to do and can get away with. Regulators still have to follow regulations. You are right that I can'…

Should you start all of your statements with a list of every project you've ever worked on? Show me an example of how it's done before you make such an exceptional request of me.

For any who are more curious about a commenters' background than their current words, my profile already links to copious resources on my work history, & writings, beyond what's typical of contributors here.

Re: Tell HN: Internet Archive is facing a Big 4 Publishers lawsuit

#215
post #214
post #208

Earlier quoted context omitted.

As I've now learned, you used to work for the Internet Archive. You should probably start your statements with that. > If you've taken no self-help measures to limit access... robots.txt was a nice self-help measure. > ... then bad actors, unobservable to you and regulators, will still be doing whatever they would like to do and can get away with. Regulators still have to follow regulations. You are right that I can'…

Should you start all of your statements with a list of every project you've ever worked on? Show me an example of how it's done before you make such an exceptional request of me. For any who are more curious about a commenters' background than their current words, my profile already links to copious resources on my work history, & writings, beyond what's typical of contributors here.

Yes, if I was commenting on a project or company that I had self-interest in (eg: reputational, monetary, etc) then I would add a disclaimer like everyone else here does.

If you're writing messaging systems and working on cryptocurrency you can figure out Algolia: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...

You didn't need me to do that.

Re: Tell HN: Internet Archive is facing a Big 4 Publishers lawsuit

#216

Earlier quoted context omitted.

This is victim blaming. In my jurisdiction, you retain copyright under any information you publish, even to the worldwide public. This means I can reasonably expect entities to collect, save, analyze and repurpose that info within reason , and without specific steps to discourage access & use. This is why there are laws such as 'fair use' and 'satire', because we wanted to extend what is considered reasonable use of…

You are arguing about copyright in a thread discussing accusations of privacy violations.

Copyright is a mechanism used to protect privacy in these situations. When you don't have copyright, you are stuck needing a court to protect your privacy. Copyright is also what is required to prove in order to get stuff taken down by IA when the content is not obviously illegal or personally identifiable information (or at least it was when I last needed to deal with it).

Re: Tell HN: Internet Archive is facing a Big 4 Publishers lawsuit

#217
post #182

Earlier quoted context omitted.

Information on a public website is public until it is taken down or the information changed. The Internet Archive removes an individuals control over when the information remains public. This is privacy. We might be caught naked, and we can't unsee what has been seen, but it is a basic human instinct to draw the curtains and contain further damage. Perfectly innocent individuals suffer because the IA rules are design…

> The Internet Archive removes an individuals control over when the information remains public. And that's a good thing in the vast majority of cases. Unless we're talking about sensitive information that was published without the consent of the person in question, all public information should remain public forever.

In my experience, it is the vast minority of cases. Most of the content of the IA is not in the public interest, now or in the future. It is crap. It is noise. It is the contents of the Internet at a point in time. Actual information is the wheat in the chaff, and why you need search engines to find it. We know this, because of the Usenet archives that are intermittently available. Almost completely useless apart from people having a giggle at how the Internet used to be, a quick browse and search for naughty words. And a few gems in the mountain of noise, in such dire need of curation people hardly know it exists and barely justifiable enough for libraries to keep it alive.

Re: Tell HN: Internet Archive is facing a Big 4 Publishers lawsuit

#218
post #201

Earlier quoted context omitted.

This is victim blaming. In my jurisdiction, you retain copyright under any information you publish, even to the worldwide public. This means I can reasonably expect entities to collect, save, analyze and repurpose that info within reason , and without specific steps to discourage access & use. This is why there are laws such as 'fair use' and 'satire', because we wanted to extend what is considered reasonable use of…

I'm talking about the unfair allegation of privacy violations, here. Note that when the Archive shares crawled content with other libraries, those other libraries often have their own legal right to collect, preserve, and make-available that data even stronger than the Archive's rights via fair use, implied-license, library privileges, and other grounds. For example, many of the Archive's partners in government libra…

Crawling and archiving everything, including personal writings, is a chilling effect. It is the same situation people are seeing with social media, where the past remains to haunt the present and none of our future leaders are using it without a mask. It was most surprising to people when some Libraries decided 'published' meant anything put on the WWW or posted to Usenet. It seemed grasp for funding and to keep relevant in an age where information was moving out of published media and into opinions virtually scrawled on a toilet door. The stuff I needed to get removed from the Australian National Library's archive is exactly the sort of stuff that shouldn't be in there, directly against the statutory rights and mission, and the sort of thing that could be pointed to when you wanted to defund the project. Because some twit thought meaningful Australian published materials meant anything under a .au top level domain, all the dross hoovered up by IA including all the stuff since removed because it is in nobodies interest or causing harm. And it was a pain in the arse.

Re: Tell HN: Internet Archive is facing a Big 4 Publishers lawsuit

#219

The centralization of the IA should've got more attention sooner. I've worried that the Wayback Machine will only remain up for another couple of years as a result of the IA's actions. It has saved me countless times in the past, but it's sadly a one-of-a-kind, fragile trove of data in the hands of an organization that didn't keep their ideals separate from reality. I feel they should be taking steps immediately to e…

The only comparable public alternative to the Wayback Machine is archive.today. But as far as I know, no one knows who operates or funds the project. It could go down anytime and no one can do anything about it.

Re: Tell HN: Internet Archive is facing a Big 4 Publishers lawsuit

#220

Slightly off topic. What is the tech stack and architecture of "archive.org". How do they manage such huge storage. How do they forecast and how much does it cost per year to manage existing storage ? I could not find any links?

I don't have the answer either. But you might be able to find some traces in their blog: http://blog.archive.org/category/technical/
Post reply on HN