Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

211–220 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#211

Earlier quoted context omitted.

the english empire once tried to mantain a monopoly over steam loom machines the americans cheated their way to competition, heck, even before that, the english empire got jumpstarted by stealing gold from the spanish (who were themselves exploiting it away from aztec and other mexican natives) I'm saying it's business as usual, but also, culture doesn't work like tangible physical widgets so we must stop letting a f…

I don't think I've heard the term "English empire". Is it an attempt by the Scottish to pretend they weren't involved?

Just like Austria's greatest historical accomplishment: convincing the world the Hitler was German

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#212
post #63
post #30

Earlier quoted context omitted.

100TB is like 6 hard drives...

> 100TB is like 6 hard drives... Discounted Seagates ? /s

You can get recertified 18TB drives, but still it's a lot of disk space. I simply don't have enough data.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#213
post #157
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

The thing is Google, meta and YouTube weren't giant entities when they did this stuff. I think it's good no one cracked down on them for copyright stuff. Now they're developing an LLM that will generate potentially trillions in value to humanity and looks like they're not exactly playing by the rules. But I prefer looser intellectual property rights anyway so Im ok with it

DRM for thee not for me.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#214
One of the largest businesses of the Internet to date has been piracy. Individual informal piracy has been the smallest component of this. By far the largest has been corporate mass-scale piracy, and LLMs are probably the largest heist to date. They've literally downloaded the sum total of all human thought and knowledge, compressed it into queryable lossy compression models (which is what LLMs are), and are selling it back to us.

Meta, with its "open weights" models, is one of the least guilty parties, since at least they've made the resulting blobs of mass piracy available to us. Same with Mistral, Deepseek, etc.

ClosedAI, Google, and others have all probably done this and more and refuse to make even the model available.

I think the way to deal with this is very simple:

If you have trained your model on works to which you do not have rights or permission, the resulting model is not copyrightable and cannot be sold. It must either be kept for research purposes only or released free of charge and in the public domain. All these models that have been trained on pirated works should become public domain.

Of course now that we have full capture of the US Federal Government I'm sure any suggestion like that would be neutralized with one bribe to Trump.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#215
post #153

Earlier quoted context omitted.

But the crime is creating something new. If laws are enforced that criminalise creation, then the world will be rather static. It seems to be a consistent direction of history's arc that the people who make it easy to create and innovate get ahead.

We don’t allow indiscriminate human experimentation in medicine. We have crimes against this, and yet we still have new medicines. Sure, it won’t be as quick if we could just use humans as test subjects from the start, but that’s an unethical line. Innovation done immorally is progress that shouldn’t have been made. The ends don’t justify the means, but I’m not an ethical nihilist. The crime is downloading and copyin…

Those medical policies have condemned thousands, possibly millions, to lives of unnecessary pain and suffering. They're more damaging than copyright.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#216
post #60

Earlier quoted context omitted.

First punish them. Then change the laws.

I bet you and my "first build the product, then worry about security" manager would get along.

My approach is same. First fire that manager. Then define security.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#218
post #185
post #58

We all like hating big corporations, especially Meta, and people seem to use this as an opportunity to advocate for punishing them. I think it's wiser to advocate for changing our IP laws.

You're conflating different problems. Big corporations are too big, they should just not exist. When you have corporations more powerful than the government of the biggest states, it's a bug, not a feature. The IP laws may need rethinking. Saying that they should disappear because big corporations are above the law doesn't help, though. First kill the big corporations, then think about fair laws. Changing the law now…

Perhaps they just did, or we are doing it - basically this should lead to abolition of copyright to any published article there is. Not sure how’d it impact open source, we’ll either have all of it open, or none at all.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#219
post #41

Considering prices for single work, this must be multi-billion dollar compensation. Take for example 675k paid for 31 songs. So 20k a song. If we estimate book to be say 10MB that would 8 million works. So I think reasonable compensation is something along 163 billion. Not even 10 years of net income. Which I think is entirely fair punishment.

Meta argues that it's fair use, and that they just downloaded, and never seeded, all the torrents.

The article is purposely conflating the downloading from the seeding statistics. Saying "just 0.008%" the size resulted in big punishments is confusing when Meta is also saying they set their client to be leechers.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#220
post #58

We all like hating big corporations, especially Meta, and people seem to use this as an opportunity to advocate for punishing them. I think it's wiser to advocate for changing our IP laws.

First we must prosecute Meta into committing suicide like was done to Aaron Swartz. After justice is served, we should change IP laws.
Post reply on HN