Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

51–60 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#51

It really makes you think about those crazy internet folks from back in the day who thought copyright law was too strict and that restricting humanity to knowledge in such a way was holding us all back for the benefit of a tiny few.

The more concerning thing is that the best thing these overpaid people could come up with was.. download the torrent, like everyone else. Here you are, billions of resources, and no one is willing to spend a part of it to at least digitize some new data? Like even Google did?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#52
post #41

Considering prices for single work, this must be multi-billion dollar compensation. Take for example 675k paid for 31 songs. So 20k a song. If we estimate book to be say 10MB that would 8 million works. So I think reasonable compensation is something along 163 billion. Not even 10 years of net income. Which I think is entirely fair punishment.

For creating a backup of library genesis. No. They should be awarded a philanthropic prize.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#54
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

Comprehensive intellectual property needs to happen for the modern (digital) era.

Basically the entire legal system needs to be retooled and rethought for computers.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#55
If you're an author with a book likely to have be hoovered up, I wonder what you'd get from the fb models if you asked "complete this in the style of [author] in [book]: [quite a long excerpt]"

If you get a direct quote then you're good with your claim, surely.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#56
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> Google itself got big by indexing other people's data without compensation

Weird framing given how much value was and is still placed on Google driving traffic to you

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#59
post #4

Earlier quoted context omitted.

I’ve got 70-80mb pirated books, I think because of the illustrations. Guess it depends on the book.

I don’t think they’re using picture heavy book for LLM training, no?

Why not ? Do you think that AI doesn't enjoy porn ? /s

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#60
post #58

We all like hating big corporations, especially Meta, and people seem to use this as an opportunity to advocate for punishing them. I think it's wiser to advocate for changing our IP laws.

First punish them. Then change the laws.
Post reply on HN