Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

31–40 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#34
post #21

Eye for an eye. Meta losses rights to 81.7 TB of IP. Transcribed into a text file

Meta already does that to themselves every year or so, deleting all internal communications. They've thrown away a huge amount of communication to source code commit reinforcement training data as a result. They do it to avoid emails making it into trials like this.

> Meta already does that to themselves every year or so, deleting all internal communications.

Aren't they obligated by law to keep all internal communication?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#37
Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the early days. The contracts with the music labels came later. GPL violations by commercial products fits the theme also.

Companies aggressively protect their own intellectual property but have no qualms about violating the IP rights of others. Companies. Individuals have no such privilege. If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#38
post #7
post #4

Earlier quoted context omitted.

I don’t think they’re using picture heavy book for LLM training, no?

For multi-modal models, why not? They would be probably some of the best data.

Sometimes the PDF of a book is big because the book's packed with important illustrations and charts - like a textbook or journal paper.

Other times a PDF of a book is big because someone scanned it and didn't have trustworthy OCR, so they figured distributing images of text at 1.5 MB per page was better than risking OCR errors.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#39
post #20

A good chance for federal prosectutors to "send a message" as they did with Aaron Swartz but I don't see things going that way.

Even after JSTOR declined to press charges in that case. Despicable. The US has dug the hole it's going down.
Post reply on HN