Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

481–490 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#481
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

VC and startups are fundamentally about disruption. You can't make an omelette without breaking a few eggs (laws). The incumbent players are not going to sit still and let things be "disrupted". A common response is to make sure the public knows about the broken eggs. I would say youtube, Google, Spotify, Uber, doordash, etc. all have made my life much better.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#482

Earlier quoted context omitted.

Critically, by torrenting they also directly distributed the copywritten material itself. That is a standalone infringement separate from any argument about trained LLMs.

They could have only leached and refrained from sharing any part of copyrighted data. If i were to commit something as risky as this, that is what i would do.

Then it would need to be determined, whether that is the case or not. Did every single machine they used have the configuration for only leeching and no seeding? The company is liable for what its employees on the job. If only one employee was also seeding ... that could be a very interesting case.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#483

"Say they hood robin, ain't that a b*, take from the poor and give to the rich." - Ice Cube. Meta will face no consequences. Say your a small publisher and you'd like a bit of compensation. If you dare sue Meta can just blacklist your books on its platforms. Even if they don't, you probably don't have the money to sue one of the biggest companies on earth. I think copyrights should be limited to 25 years after first…

can people vote with their feet, and leave the platform ?

but the masses are addicted to the slop that meta feeds them.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#484
post #404

Earlier quoted context omitted.

What's interesting is that in many cities now, Uber and Lyft are in fact more expensive than taxis. And the experience is equally mediocre. The pendulum has swung back the other way. The only thing they have going for them now is the app based convenience, which is eroding as more "yellow cab" type traditional taxis band together and get set up with their own sort of city-specific app.

I remember calling a taxi 3 hours before my flight to get to SFO. After an hour and four different phone calls to the taxi company, I took BART and barely made it before the counter closed. The feedback system incentivizes drivers and riders to behave.

This is getting off-topic, but I am curious, why didn't you go with BART in the first place? If you had an hour to call the taxi company and still arrive in time, presumably, you had more than enough time.

I know there are reasons for not going with public transport, but preferring to take a taxi/uber when a train line can get you there in time maybe has more to say about public transport than about taxis. Well functioning rail is typically one of the most effective and reliable way of getting to an airport, and often much cheaper than taxis.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#485
post #362

Earlier quoted context omitted.

Critically, by torrenting they also directly distributed the copywritten material itself. That is a standalone infringement separate from any argument about trained LLMs.

And punishing them in the normal manner will be an incredibly small slap on the wrist, and do absolutely nothing to help us find out what will play out in court regarding a fair-use defense on training AI with copyrighted material.

Isn't there a "fruit of the poisoned tree" kind of thing? Sounds to me quite similar to the situation where you would murder your parent and get to keep the inheritance, even if you are convicted of murder. Inheriting stuff isn't illegal, yet, I think most jurisdictions would not allow you to keep it in this case.

There should be a problem with stuff obtained through illegal means, even if having that stuff is in principle legal. In this case, copyrighted material.

Obviously they would argue that having the data is only a consequence of the download part, and that part is legal. What I see is that these situations are always complicated, and if you're rich enough, you get to litigate the complications and come out with a slap on the wrist or maybe even clean hands, while if you are an ordinary citizen, you can't afford to delve into the complexities and get punished.

These days I'm starting to give up on the whole concept of the legal system being fair. They're not even pretending anymore.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#486

Earlier quoted context omitted.

good distinction IMO there's a hack about this, authors can claim that they allow for public use unless it's used for training LLMs. And all of training work would fall under 2 because they would be used against the copyright.

I think they would need to have some explicit contract every time they want to sell the book then, though. I don’t think I am bound by some random terms someone writes into a book I’m buying. Those probably are only binding if a reasonable person would notice them before sale.

If you arrive at the point of being able to buy that book, it means it has passed the publisher's hands and I would think, that the publisher was OK with those terms then, and limiting the usage of the text may in fact be effective. If it was self-published, then even more so.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#487
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

Aaron committed suicide and FBI going after him was meant more as a lesson to the other kids at MIT than anything.

MegaUpload did the same, kim dotcom got raided in his sleep by FBI in New Zealand! So no I don't buy your reductionist argument, there are forces at play that allow companies with founders with the likes of Google to get away with it but not others.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#488
post #166

Earlier quoted context omitted.

Airbnb and Uber have showed us that laws matter only to the extent that the political will to enforce them exists. Throw enough lawyers and lobbying money at the problem and the laws can simply be re-written to be friendlier to your business model.

The hotel and taxi industry were legit terrible before those two disrupted them. Laws are ment to be broken. Especially in cronist systems where incumbents write the laws.

The level of casual criminality in this industry is astounding sometimes.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#489
post #58

We all like hating big corporations, especially Meta, and people seem to use this as an opportunity to advocate for punishing them. I think it's wiser to advocate for changing our IP laws.

The point is about the hypocrisy and double-standards evinced by this behavior.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#490
post #329

Earlier quoted context omitted.

Everyone on here is smart enough. Just do not participate and save your money. Do not pay for digital goods. If Netflix raises their prices, it doesn't matter because there is a torrent of all of their shows. If Spotify raises their prices, it doesn't matter because your favorite artist has their entire library in a torrent. If some game company ask you to pay real life prices for a digital costume, find the crack on…

I just can't get behind the sentiment that the unethical behavior by big companies means I get to access all the content I want for free.

Imo it’s not about you accessing things you want for free. If your family purchased a disc copy of the goonies before you were born and you watched it as a kid, your accessing of that content you wanted for free has no moral bearing. The core question is what impact does your consumption have, and I don’t think that participating in the streaming landscape is making things any better for anyone but their ceos.
Post reply on HN