Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

441–450 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#441
post #433

Earlier quoted context omitted.

The right to read and study you have by default . It's getting your hands on a book that has legal caveats attached.

Yes, but getting your hands on the material isn't a very interesting legal question IMO. Whether you can train your LLM on it is a very interesting question. I've personally never been in favor of punishing people for downloading (or seeding) things.

One which buying books for your LLM doesn't answer either. In analogy to humans, you might as well give your LLM a library card.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#442

Earlier quoted context omitted.

I rooted for Uber to smash the Taxi cartels. Let us not forget that Taxi Cartels were also insidious beasts. Taxi drivers abused their walled garden with their price gouging by taking longer routes, refusal to take a credit card, and extremely poorly maintained fleets of vehicles. I have had mostly good experiences with Uber, whereas I had experiences that mostly bordered on general condescension toward me whenever I…

What's interesting is that in many cities now, Uber and Lyft are in fact more expensive than taxis. And the experience is equally mediocre. The pendulum has swung back the other way. The only thing they have going for them now is the app based convenience, which is eroding as more "yellow cab" type traditional taxis band together and get set up with their own sort of city-specific app.

It does always seem like a race to the bottom.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#443
post #351

Earlier quoted context omitted.

Or if you don't agree with the price, do not buy it. You are NOT entitled to entertain yourself in any way you want. (unless it's funded by taxes etc, in which case... okay, it's open to discussion.) Look, let's be honest - what gives you or others the right to steal from others?

That's what I do, personally. But you call it "stealing," others call it "copying." Stealing takes, from someone, something they own.

There's such a mass of possible works that it hardly constrains someone that if you could cast a magic spell preventing someone from distributing or accessing your particular work and then burned it, your spell would have essentially no effect-- no one would notice it and no one would be harmed.

As long as discussion of a work that has published is not impeded, the public is not harmed even by these 50-years after life copyrights other than by that they are accumulated by certain companies who themselves become problems.

When someone decides to use someone's work without compensation he is, even though he is not deprived of the work itself, still robbed. But it's not a theft of goods, it's theft of service. The copyright infringer isn't the guy who steals your phone, it's the guy who even you have done some work for but who refuses to pay.

With this view you can also believe, without hypocrisy, that what the LLM firms are doing is wrong while what Schwartz did was not, since the authors in question weren't deprived of any royalties or payments due to them due to due to the publishing model for scientific works.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#444

I don't understand why it's even a question that Meta trained their LLM on copyrighted material. They say so in their paper! Quoting from their LLaMMa paper [Touvron et al., 2023]: > We include two book corpora in our training dataset: the Gutenberg Project, [...], and the Books3 section of ThePile (Gao et al., 2020), a publicly available dataset for training large language models. Following that reference: > Books3…

Critically, by torrenting they also directly distributed the copywritten material itself. That is a standalone infringement separate from any argument about trained LLMs.

They could have only leached and refrained from sharing any part of copyrighted data. If i were to commit something as risky as this, that is what i would do.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#445
post #193

Earlier quoted context omitted.

i know of a company that poisoned an entire town! thats terrorism if done by an individual. the company still exists, just paid a settlement and carried on...

I agree with your point, but will split hairs on using the word "terrorism". I think that should be reserved for people that commit atrocities for some political aim. I'm fairly sure the company in question (I assume Union Carbide) did not poison the town to advance a political agenda.

Money is the ultimate political agenda.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#446
post #263

Earlier quoted context omitted.

old books? i can imagine the shit/hallucinated-like generative AI we would have if the training weight was restricted to public domain stuff... i think when chatGPT was around version 2 or 3, i had extracted almost 2 pages (without any alteration from the original) with questions that considered the author from this book here, https://www.amazon.com/Loneliness-Human-Nature-Social-Connec... now it's up to you to think…

I find this such a strange remark on this front. You got less than 1% of a book... from an author who has passed away... who wrote on a research topic that was funded by an institution that takes in hundreds of millions of dollars in federal grants each year... I'm not an author (although I do generate almost exclusively IP for a living) and I think this is about as weak a form of this argument as you possibly make.…

I think the key is to think through the incentives for future authors.

As a thought experiment, say that the idea someday becomes mainstream that there is no reason to read any book or research publication because you can just ask an AI to describe and quote at length from the contents of anything you might want to read. In such a future, I think it's reasonable to predict that there would be less incentive to publish and thus less people publishing things.

In that case, I would argue the "hurt" is primarily to society as a whole, and also to people who might have otherwise enjoyed a career in writing.

Having said that, I don't think we're particularly close to living in that future. For one thing I'd say that the ability to receive compensation from holding a copyright doesn't seem to be the most important incentive for people to create things (written or otherwise), though it is for some people. But mostly, I just don't think this idea of chatting with an AI instead of reading things is very mainstream, maybe at least in part because it isn't very easy to get them to quote at length. What I don't know is whether this is likely to change or how quickly.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#447
post #263

Earlier quoted context omitted.

old books? i can imagine the shit/hallucinated-like generative AI we would have if the training weight was restricted to public domain stuff... i think when chatGPT was around version 2 or 3, i had extracted almost 2 pages (without any alteration from the original) with questions that considered the author from this book here, https://www.amazon.com/Loneliness-Human-Nature-Social-Connec... now it's up to you to think…

I find this such a strange remark on this front. You got less than 1% of a book... from an author who has passed away... who wrote on a research topic that was funded by an institution that takes in hundreds of millions of dollars in federal grants each year... I'm not an author (although I do generate almost exclusively IP for a living) and I think this is about as weak a form of this argument as you possibly make.…

It isn't that someone was hurt. We have one private entity gaining power by centralizing knowledge (which they never contributed to) and making people pay for regurgitating the distilled knowledge, for profit.

Few entities can do that (I can't).

Most people are forced to work for companies that sell their work to the higher bidder (which are the very entities mentioned above), or ask them to use AI (under the condition that such work is accessible to the AI entities).

It's obviously a vicious circle, if people can't oppose their work to be ingested and repackaged by a few AI giants.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#449
post #393
post #342

Earlier quoted context omitted.

> Say your rents it's going to be late, you can pick up 20 or 30 hours of Uber this month to make it happen that sounds so incredibly dystopian, not sure if that was the intention :(

It's super dystopian and it creates bad incentives (the harder the underclass is squeezed the better a product it is for the middle class), but I have to agree that gig work is often a lifeline for poor people. I consider it similar to access to unsecured credit that way - it's easy to feel like "wow this industry is scamming these people it should be illegal" but people without any other backstop will probably need…

maintaining a precarious class benefits those in power

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#450
post #394

Earlier quoted context omitted.

The answer is to censor the model output, not the training input. A dumb filter using 20 year old technology can easily stop LLM's from verbatim copyright output.

What if the model simply substitutes synonyms here and there without changing the spirit of the material? (This might not work for poetry, obviously.) It is not such a simple matter.

It's pretty simple, you are absolutely allowed to do that, and it's been done forever.

Imagine having the copyright claim to "Person's family member is killed so they go and get revenge".

Post reply on HN