Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

1–10 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#4

ebooks are a 1-2 mb each max. 81.7 TB are a lot of books, like 42-85 million books.

I’ve got 70-80mb pirated books, I think because of the illustrations. Guess it depends on the book.

I don’t think they’re using picture heavy book for LLM training, no?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#6
post #4

Earlier quoted context omitted.

I’ve got 70-80mb pirated books, I think because of the illustrations. Guess it depends on the book.

I don’t think they’re using picture heavy book for LLM training, no?

Yes they do, there's multimodal models.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#7
post #4

Earlier quoted context omitted.

I’ve got 70-80mb pirated books, I think because of the illustrations. Guess it depends on the book.

I don’t think they’re using picture heavy book for LLM training, no?

For multi-modal models, why not? They would be probably some of the best data.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#9
post #4

Earlier quoted context omitted.

I’ve got 70-80mb pirated books, I think because of the illustrations. Guess it depends on the book.

I don’t think they’re using picture heavy book for LLM training, no?

Presumably they didn't create the torrent
Post reply on HN