[flagged]
Meta torrented & seeded 81.7 TB dataset containing copyrighted data
31–40 of 981 posts
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#32Before I decided my opinion on this I need to know their ratio.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#33[flagged]
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#34Eye for an eye. Meta losses rights to 81.7 TB of IP. Transcribed into a text file
Meta already does that to themselves every year or so, deleting all internal communications. They've thrown away a huge amount of communication to source code commit reinforcement training data as a result. They do it to avoid emails making it into trials like this.
Aren't they obligated by law to keep all internal communication?
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#35Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#36Remembering Aaron Swartz in this moment
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#37Companies aggressively protect their own intellectual property but have no qualms about violating the IP rights of others. Companies. Individuals have no such privilege. If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#38Earlier quoted context omitted.
I don’t think they’re using picture heavy book for LLM training, no?
For multi-modal models, why not? They would be probably some of the best data.
Other times a PDF of a book is big because someone scanned it and didn't have trustworthy OCR, so they figured distributing images of text at 1.5 MB per page was better than risking OCR errors.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#39A good chance for federal prosectutors to "send a message" as they did with Aaron Swartz but I don't see things going that way.