Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

191–200 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#191

That's horrible! Magnet anyone?

Anna's Archive: https://annas-archive.org

specifically https://annas-archive.se/torrents - this is a meta-project which aggregates illegal copyrighted material from other illegal projects. You absolutely should not download any material this page links to, although you can use it for the purpose of researching about shadow libraries.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#192
post #21

Earlier quoted context omitted.

Meta already does that to themselves every year or so, deleting all internal communications. They've thrown away a huge amount of communication to source code commit reinforcement training data as a result. They do it to avoid emails making it into trials like this.

> Meta already does that to themselves every year or so, deleting all internal communications. Aren't they obligated by law to keep all internal communication?

Companies often don't do what they're obligated to. As long as they can keep plausible deniability.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#193
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

i know of a company that poisoned an entire town! thats terrorism if done by an individual. the company still exists, just paid a settlement and carried on...

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#194

It really makes you think about those crazy internet folks from back in the day who thought copyright law was too strict and that restricting humanity to knowledge in such a way was holding us all back for the benefit of a tiny few.

I'm all for chopping up copyright law. But until we do so, companies like Meta need to be treated just like everyone else.

That means lawsuits, prison sentences, and millions in fines. And that's just the piracy part, there's also the lying/fraud part.

Interestingly, a Dutch LLM project was sent a cease and desist after the local copyright lobby caught wind of it being trained on a bunch of pirated eBooks. The case unfortunately wasn't fought out in court, because I would be very interested to see if this could make that copyright lobby take down ChatGPT and the other AI companies for doing the same.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#195
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> Spotify's music library was also pirated in the early days.

I want to know more, please enlighten me (anyone who knows). I read the book "The Spotify Play" and it made it seem like the pirated music was an internal-only thing and not something available to customers. Is that true?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#196
post #41

Considering prices for single work, this must be multi-billion dollar compensation. Take for example 675k paid for 31 songs. So 20k a song. If we estimate book to be say 10MB that would 8 million works. So I think reasonable compensation is something along 163 billion. Not even 10 years of net income. Which I think is entirely fair punishment.

Beyond the absurdity of those amounts, the funny thing is that the authors wouldn’t ever see a dime of that money. Not in the music case, not in this one either. Fairness?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#197
post #58

We all like hating big corporations, especially Meta, and people seem to use this as an opportunity to advocate for punishing them. I think it's wiser to advocate for changing our IP laws.

Big corporations don't have morale or ethics. They'll break any laws as long as it's profitable. There's no point complaining about Meta or Zuck. Meta does what it's designed to do. If people aren't happy, they should vote for more regulations.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#198

Earlier quoted context omitted.

Comprehensive intellectual property needs to happen for the modern (digital) era. Basically the entire legal system needs to be retooled and rethought for computers.

No we just need to enforce the existing laws. And the legal system is for humans not computers.

Plenty of the existing laws are insane and indefensible. Copyright duration of life of the author plus 70 years? Patents on videogame mechanics?

We need to both reform the laws and enforce them. Otherwise...

>The law, in its majestic equality, forbids the rich and poor alike to sleep under bridges, to beg in the streets, and to steal bread.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#199
post #123

Libgen is a civilizational project that should be endorsed, not prosecuted. I hope one day people will look at it and think how stupid we were today to shun the largest collection of literary works in human history.

I think you’re overstating its importance. The internet already makes it possible to order almost any book in existence and have it arrive at your doorstep within a week or so, or often on your ebook reader instantly. And your local library probably participates in an interlibrary loan system that lets you request any book held by any library in the country for free . LibGen gives you access to a much smaller body of…

And what about the other billions of people on the planet that don't even have a library, let alone a doorstep to receive a first world delivery service.
Post reply on HN