Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

21–30 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#21

Eye for an eye. Meta losses rights to 81.7 TB of IP. Transcribed into a text file

Meta already does that to themselves every year or so, deleting all internal communications.

They've thrown away a huge amount of communication to source code commit reinforcement training data as a result. They do it to avoid emails making it into trials like this.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#22
post #8

Really curious what the judges are going to do here. Horse has functionally bolted on this already I’m guessing slap on wrist despite courts going after individual for a couple of movies torrented pretty hard

[flagged]

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#25

ebooks are a 1-2 mb each max. 81.7 TB are a lot of books, like 42-85 million books.

The article says they got datasets from Anna's Archive. It was most likely the scihub/libgen torrent which is 96.0 TB right now and contains 92,872,581 files. That's about 1 megabyte per file.

https://annas-archive.org/datasets

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#26

So if I torrented and seeded, I would be doing it for my own entertainment, not commercially. I expect big copy-write holders to come after myself. If Meta does it - I guess they have better lawyers ? Could make interesting case law.

> Could make interesting case law.

Yeah, to perpetuate this system where only those who can afford lawyers get to benefit

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#29
post #4

Earlier quoted context omitted.

I’ve got 70-80mb pirated books, I think because of the illustrations. Guess it depends on the book.

I don’t think they’re using picture heavy book for LLM training, no?

Just because the LLMs are trained on text doesn't mean that images we're a part of what they downloaded.

You clean up the data after you acquire it, not before.

Post reply on HN