Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

11–20 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#13
post #4

Earlier quoted context omitted.

I’ve got 70-80mb pirated books, I think because of the illustrations. Guess it depends on the book.

I don’t think they’re using picture heavy book for LLM training, no?

Even if they didn't use the illustration(which isn't clear given multimodal models), they'd still make use the text in the books.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#15
post #4

Earlier quoted context omitted.

I don’t think they’re using picture heavy book for LLM training, no?

Presumably they didn't create the torrent

Whoever created it has a lot of spare hard disk space.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#18
post #4

Earlier quoted context omitted.

I’ve got 70-80mb pirated books, I think because of the illustrations. Guess it depends on the book.

I don’t think they’re using picture heavy book for LLM training, no?

I don't think they need to be selective. It's not like Meta can run out of storage.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#19
post #11

So according to some AI, the damages awarded per infringed work is ~$750 minimum in the US. 80TB of books, each let's say 10MB on average, would be 8 million works. So Meta should pay 6 billion USD for their copyright infringement?

Nice calculation, that’s actually quite doable for them, they have already been paying similar fines for a while.
Post reply on HN