That's horrible! Magnet anyone?
Anna's Archive: https://annas-archive.org
Meta torrented & seeded 81.7 TB dataset containing copyrighted data
191–200 of 981 posts
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#192Earlier quoted context omitted.
Meta already does that to themselves every year or so, deleting all internal communications. They've thrown away a huge amount of communication to source code commit reinforcement training data as a result. They do it to avoid emails making it into trials like this.
> Meta already does that to themselves every year or so, deleting all internal communications. Aren't they obligated by law to keep all internal communication?
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#193Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#194It really makes you think about those crazy internet folks from back in the day who thought copyright law was too strict and that restricting humanity to knowledge in such a way was holding us all back for the benefit of a tiny few.
That means lawsuits, prison sentences, and millions in fines. And that's just the piracy part, there's also the lying/fraud part.
Interestingly, a Dutch LLM project was sent a cease and desist after the local copyright lobby caught wind of it being trained on a bunch of pirated eBooks. The case unfortunately wasn't fought out in court, because I would be very interested to see if this could make that copyright lobby take down ChatGPT and the other AI companies for doing the same.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#195Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…
I want to know more, please enlighten me (anyone who knows). I read the book "The Spotify Play" and it made it seem like the pirated music was an internal-only thing and not something available to customers. Is that true?
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#196Considering prices for single work, this must be multi-billion dollar compensation. Take for example 675k paid for 31 songs. So 20k a song. If we estimate book to be say 10MB that would 8 million works. So I think reasonable compensation is something along 163 billion. Not even 10 years of net income. Which I think is entirely fair punishment.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#197We all like hating big corporations, especially Meta, and people seem to use this as an opportunity to advocate for punishing them. I think it's wiser to advocate for changing our IP laws.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#198Earlier quoted context omitted.
Comprehensive intellectual property needs to happen for the modern (digital) era. Basically the entire legal system needs to be retooled and rethought for computers.
No we just need to enforce the existing laws. And the legal system is for humans not computers.
We need to both reform the laws and enforce them. Otherwise...
>The law, in its majestic equality, forbids the rich and poor alike to sleep under bridges, to beg in the streets, and to steal bread.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#199Libgen is a civilizational project that should be endorsed, not prosecuted. I hope one day people will look at it and think how stupid we were today to shun the largest collection of literary works in human history.
I think you’re overstating its importance. The internet already makes it possible to order almost any book in existence and have it arrive at your doorstep within a week or so, or often on your ebook reader instantly. And your local library probably participates in an interlibrary loan system that lets you request any book held by any library in the country for free . LibGen gives you access to a much smaller body of…