Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

41–50 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#41
Considering prices for single work, this must be multi-billion dollar compensation.

Take for example 675k paid for 31 songs. So 20k a song. If we estimate book to be say 10MB that would 8 million works. So I think reasonable compensation is something along 163 billion. Not even 10 years of net income. Which I think is entirely fair punishment.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#43
post #21

Eye for an eye. Meta losses rights to 81.7 TB of IP. Transcribed into a text file

Meta already does that to themselves every year or so, deleting all internal communications. They've thrown away a huge amount of communication to source code commit reinforcement training data as a result. They do it to avoid emails making it into trials like this.

No large company will ever consider training a public LLM on all their internal communications.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#44
post #21

Earlier quoted context omitted.

Meta already does that to themselves every year or so, deleting all internal communications. They've thrown away a huge amount of communication to source code commit reinforcement training data as a result. They do it to avoid emails making it into trials like this.

> Meta already does that to themselves every year or so, deleting all internal communications. Aren't they obligated by law to keep all internal communication?

Yes, they are. But I can imagine the fine/impact for this being much, much lower than the consequences of all their nefarious communication being used in trials.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#47

So if I torrented and seeded, I would be doing it for my own entertainment, not commercially. I expect big copy-write holders to come after myself. If Meta does it - I guess they have better lawyers ? Could make interesting case law.

> Could make interesting case law. Yeah, to perpetuate this system where only those who can afford lawyers get to benefit

Since it’s case law, everyone would benefit from the precedent

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#48
post #11

So according to some AI, the damages awarded per infringed work is ~$750 minimum in the US. 80TB of books, each let's say 10MB on average, would be 8 million works. So Meta should pay 6 billion USD for their copyright infringement?

Minimum doesn't cover willful copyright infractions, for which maximum penalty is $150K per work. That comes out to quite a different number.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#49
post #21

Earlier quoted context omitted.

Meta already does that to themselves every year or so, deleting all internal communications. They've thrown away a huge amount of communication to source code commit reinforcement training data as a result. They do it to avoid emails making it into trials like this.

> Meta already does that to themselves every year or so, deleting all internal communications. Aren't they obligated by law to keep all internal communication?

When there is a specific order after proceeding starts, but not before. There can sometimes be other orders as part of govt settlements like Google was recently accused of violating.

You may be thinking of certain financial institutions where it is a hard requirement, and maybe there are some other regulated industries too that have it.

Post reply on HN