Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

61–70 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#62
post #56
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> Google itself got big by indexing other people's data without compensation Weird framing given how much value was and is still placed on Google driving traffic to you

Even before the LLM-craze Google was showing their Answers box or whatever it was called at the top of the results that told you the answer (sometimes) so that you didn’t have to visit any website.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#65
Best way to "punish" Meta is to slash the Gordian knot and abolish copyright. Level the playing field, incrementally, for everyone else who isn't a trillion-dollar corporation.

The alternative is a futile legalistic attack against a monopoly entity too powerful to be meaningfully punished. That won't accomplish anything useful. It would, rather, help cement this status quo, where copyright infringement is selectively legal or illegal, for different entities at the same time; and companies like Meta thrive arbitraging that difference. You can't defeat Meta—but you can help dig them a moat.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#66
post #56
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> Google itself got big by indexing other people's data without compensation Weird framing given how much value was and is still placed on Google driving traffic to you

For Google's case the order was reversed.

Google used to send customers to your site. Now they try to show you the information on their site so that the customer doesn't need to go to your site.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#67
post #58

We all like hating big corporations, especially Meta, and people seem to use this as an opportunity to advocate for punishing them. I think it's wiser to advocate for changing our IP laws.

I think most of the public is probably in favor of stronger IP laws now that big corps are threatening to make them jobless with IP-disrespecting AIs

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#68
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

the english empire once tried to mantain a monopoly over steam loom machines

the americans cheated their way to competition,

heck, even before that, the english empire got jumpstarted by stealing gold from the spanish (who were themselves exploiting it away from aztec and other mexican natives)

I'm saying it's business as usual, but also, culture doesn't work like tangible physical widgets so we must stop letting a few steal this boon of digital copying by means of silly ideas like DRM, copyright, patents. all means to cause scarcity

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#69
post #8

Really curious what the judges are going to do here. Horse has functionally bolted on this already I’m guessing slap on wrist despite courts going after individual for a couple of movies torrented pretty hard

Is there any other possible outcome than a fine? That too one which will not really affect Meta's overall earnings

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#70

That's horrible! Magnet anyone?

Library genesis

Weird shenanigans are happening in libgen at the moment; better go through Anna's Archive to look for the items you want, it will link you to the corresponding mirrors more reliably.

At least this has been the recent experience of a friend who used libgen and anna's archive to download legal, public domain works!

Post reply on HN