Earlier quoted context omitted.
the english empire once tried to mantain a monopoly over steam loom machines the americans cheated their way to competition, heck, even before that, the english empire got jumpstarted by stealing gold from the spanish (who were themselves exploiting it away from aztec and other mexican natives) I'm saying it's business as usual, but also, culture doesn't work like tangible physical widgets so we must stop letting a f…
I don't think I've heard the term "English empire". Is it an attempt by the Scottish to pretend they weren't involved?
Meta torrented & seeded 81.7 TB dataset containing copyrighted data
211–220 of 981 posts
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#212Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#213Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…
The thing is Google, meta and YouTube weren't giant entities when they did this stuff. I think it's good no one cracked down on them for copyright stuff. Now they're developing an LLM that will generate potentially trillions in value to humanity and looks like they're not exactly playing by the rules. But I prefer looser intellectual property rights anyway so Im ok with it
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#214Meta, with its "open weights" models, is one of the least guilty parties, since at least they've made the resulting blobs of mass piracy available to us. Same with Mistral, Deepseek, etc.
ClosedAI, Google, and others have all probably done this and more and refuse to make even the model available.
I think the way to deal with this is very simple:
If you have trained your model on works to which you do not have rights or permission, the resulting model is not copyrightable and cannot be sold. It must either be kept for research purposes only or released free of charge and in the public domain. All these models that have been trained on pirated works should become public domain.
Of course now that we have full capture of the US Federal Government I'm sure any suggestion like that would be neutralized with one bribe to Trump.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#215Earlier quoted context omitted.
But the crime is creating something new. If laws are enforced that criminalise creation, then the world will be rather static. It seems to be a consistent direction of history's arc that the people who make it easy to create and innovate get ahead.
We don’t allow indiscriminate human experimentation in medicine. We have crimes against this, and yet we still have new medicines. Sure, it won’t be as quick if we could just use humans as test subjects from the start, but that’s an unethical line. Innovation done immorally is progress that shouldn’t have been made. The ends don’t justify the means, but I’m not an ethical nihilist. The crime is downloading and copyin…
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#216Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#217Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#218We all like hating big corporations, especially Meta, and people seem to use this as an opportunity to advocate for punishing them. I think it's wiser to advocate for changing our IP laws.
You're conflating different problems. Big corporations are too big, they should just not exist. When you have corporations more powerful than the government of the biggest states, it's a bug, not a feature. The IP laws may need rethinking. Saying that they should disappear because big corporations are above the law doesn't help, though. First kill the big corporations, then think about fair laws. Changing the law now…
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#219Considering prices for single work, this must be multi-billion dollar compensation. Take for example 675k paid for 31 songs. So 20k a song. If we estimate book to be say 10MB that would 8 million works. So I think reasonable compensation is something along 163 billion. Not even 10 years of net income. Which I think is entirely fair punishment.
Meta argues that it's fair use, and that they just downloaded, and never seeded, all the torrents.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#220We all like hating big corporations, especially Meta, and people seem to use this as an opportunity to advocate for punishing them. I think it's wiser to advocate for changing our IP laws.