Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

461–470 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#462

I strongly urge people to read Thomas Babington Macaulay's speeches on copyright, its aims, terms, and hazards. Very well reasoned and explained. In particular, people often cited the case of authors who had died leaving a family in destitution, and claimed that copyright extension would be a fair way of preventing this, but in most cases the remaining family had never held the copyright; the author had initally sold…

> in most cases the remaining family had never held the copyright; the author had initally sold the reproduction rights to a publisher He was able to sell it because it is something valuable, exactly because of the copyright protections. Regardless of whether author sells the rights or not, he and his family would equally be better off with copyright.

Why does this argument remind me so much of those of slavery apologist arguments?

copyright as written serves the interests of publishers who don't create valuable works more than the creators of the work...

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#463

Earlier quoted context omitted.

I rooted for Uber to smash the Taxi cartels. Let us not forget that Taxi Cartels were also insidious beasts. Taxi drivers abused their walled garden with their price gouging by taking longer routes, refusal to take a credit card, and extremely poorly maintained fleets of vehicles. I have had mostly good experiences with Uber, whereas I had experiences that mostly bordered on general condescension toward me whenever I…

What's interesting is that in many cities now, Uber and Lyft are in fact more expensive than taxis. And the experience is equally mediocre. The pendulum has swung back the other way. The only thing they have going for them now is the app based convenience, which is eroding as more "yellow cab" type traditional taxis band together and get set up with their own sort of city-specific app.

> What's interesting is that in many cities now, Uber and Lyft are in fact more expensive than taxis.

Sure, agreed.

> And the experience is equally mediocre.

Absolutely not. I regret using a taxi nearly every time I opt for the cheaper option. It's only the "better" choice if you happen to be standing right in front of one. This experience is nearly universal no matter where I travel.

I think people really forget how utterly terrible Taxis were pre-Uber. I have no idea about competing apps these days, maybe they are similar to Uber, but the typical Taxi experience is nearly as awful as it's always been at least in the US.

Uber/Lyft certainly has gotten worse - but at least I can fairly reliably get a car when I need it with reasonable reliability. The rest of the "soft" product or pricing I really care far, far, less about than that simple fact.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#464

Earlier quoted context omitted.

There are two different things when it comes to discussing training LLM's on "copyright" protected data, and I almost never see people differentiate. 1.) Training on copyright that is publicly available. You write a poem and publish it online for the world to read. That is your IP, no one else can take it an sell it, but they are free to read and be inspired by it. The legalitly of training on this is in the courts,…

good distinction IMO there's a hack about this, authors can claim that they allow for public use unless it's used for training LLMs. And all of training work would fall under 2 because they would be used against the copyright.

I think they would need to have some explicit contract every time they want to sell the book then, though. I don’t think I am bound by some random terms someone writes into a book I’m buying. Those probably are only binding if a reasonable person would notice them before sale.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#465
post #166

Earlier quoted context omitted.

Airbnb and Uber have showed us that laws matter only to the extent that the political will to enforce them exists. Throw enough lawyers and lobbying money at the problem and the laws can simply be re-written to be friendlier to your business model.

The reason there was no political will to punish Airbnb and Uber for violating the law was that initially they were subsidized with VC money and so were able to undercut traditional hotels and taxis on price. In the world of tradable goods, pricing below cost with the intent of putting competition out of business so you can raise prices later is known as "dumping" and is itself illegal.

Taxis and hotels suck compared to Airbnb and uber even at the same price, so I find it hard to be upset.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#466

It really makes you think about those crazy internet folks from back in the day who thought copyright law was too strict and that restricting humanity to knowledge in such a way was holding us all back for the benefit of a tiny few.

Probably the single biggest thing I learned growing up is that you can safely live by "Everyone is in it for themselves". It's incredibly rare to find people who hold ideals that are detrimental to their own life.

Hence why I became obsessed with just about the only Philosopher who engaged with this idea seriously: https://en.wikipedia.org/wiki/The_Ego_and_Its_Own

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#467
post #376

Earlier quoted context omitted.

A model is not a backup.

Then why are we mad about the copyright stuff?

Both things can be true:

* an AI model is not a backup of the contents of all of the books in the sense that it would preserve their contents or similar such it might e.g. be useful for future generations

* Meta has (allegedly) been unfairly benefiting / profiting off of the copyrighted work of others by illegally reproducing copies of their work. Not just in the AI model sense[1], but actually (allegedly) downloading them directly from pirate repositories in a way that isn't straightforwardly fair use and even uploading some amount of this pirate data in return.

I feel like the parent commenter may have been making the typical argument for preservation of copyrighted materials, and I'm amenable to it... when it's regular people or non-profits doing that work, in a way that doesn't allow them to benefit unfairly or profit off of the hard work of others (or would be connected to such a process in some way).

Plaintiffs allege that Meta didn't just do all this, but also talked about how wrong it was and how to mitigate the seeding so they might upload as little as possible. So no matter how you slice it they allegedly 1) knew they were doing something at least a little bit wrong and 2) took steps to prevent the process that might otherwise have preserved the copied materials for the public interest.

And I feel like you probably knew all this, but maybe I'm missing something.

1: the typical argument wherein the model wouldn't exist without the ingested data, a lot of it is still in there, it is of course a derivative work and the question is really how derivative is it and what part of the work can they claim is their own contribution

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#468
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

[dead]

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#469
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life.

I’m opposed to copyright and pro-aaronsw, but the state did not kill him.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#470
post #345
post #185

Earlier quoted context omitted.

You're conflating different problems. Big corporations are too big, they should just not exist. When you have corporations more powerful than the government of the biggest states, it's a bug, not a feature. The IP laws may need rethinking. Saying that they should disappear because big corporations are above the law doesn't help, though. First kill the big corporations, then think about fair laws. Changing the law now…

How do you suggest making them smaller? For instance, what if google was still just serving search results w/ ads, and they never expanded that. How would you make them smaller?

Then they'd already be smaller, so there's no reason to make them smaller. Or am I misunderstanding your question?
Post reply on HN