Live data from Hacker News

A federal judge sides with Anthropic in lawsuit over training AI on books

techcrunch.com

141–150 of 222 posts

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#141

Earlier quoted context omitted.

> get fair compensation for the amount of work This is a bit distorted. This is a better summary: The primary purpose of copyright is to induce and reward authors to create new works and to make those works available to the public to enjoy. The ultimate purpose is to foster the creation of new works that the public can read and written culture can thrive. The means to achieve this is by ensuring that the authors of s…

As you point out, people make rules ("laws") which benefit them . I care about fairness and justice though, even if I am a minority. Fundamentally, fair compensation is based on the amount of work put in (obviously taking skill/competence into account but the differences between people in most disciplines probably don't span a single order of magnitude, let alone several). The ultimate goal should be to prevent peopl…

>Fundamentally, fair compensation is based on the amount of work put in.

I think there is a problem with your initial position. Nobody is entitled to compensation for simply working on something. You have to work on things that people need or want. There is no such thing "fair compensation".

It is "unfair" to take the work of somebody else and sell it as your own. (I don't think the LLMs are doing this.)

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#142
post #118

> “We will have a trial on the pirated copies used to create Anthropic’s central library and the resulting damages,” Judge Alsup wrote in the decision. “That Anthropic later bought a copy of a book it earlier stole off the internet will not absolve it of liability for theft but it may affect the extent of statutory damages.” I'm not sure why this alone is considered a separate issue from training the AI with books. B…

The ruling suggests that "pirating a book that could have been bought at a bookstore" for the sake of "writing a book review" "is inherently, irredeemably infringing". Which suggests that, at least in the judge's opinion, 'fair use rights' do exist in a sense, but it's about when you read the book, not when you publish. But that's not settled precedent. Meta is currently arguing the opposite in Kadrey v. Meta: they'r…

[dead]

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#144

Good. Reading books is legal. If I own a book and feed it to a program I wrote (and I have done exactly that), it is also legal. There is zero reason this should be any different with an AI.

I've co-authored a book that a lot of the models seem to know about. The models consistently get the names of the authors incorrect and quote the material with errors. If the canonical representation of our work is now embedded within AI models, don't we deserve to have it quoted and represented correctly and fairly? If you asked a human who had read the book, I think there is a fair chance they would likely give you the reference to the source material.

I do concede that the book does contain a distillation of material that is also available from other sources, but it also contained a lot of personal experience. That aspect does seem to be lost in this new representation.

I am not saying that letting AI models read the material is wrong, but the hubris in the way models answer questions is annoying.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#145

Earlier quoted context omitted.

Yep, broadly capable open models are on track for annihilation. The cost of legally obtaining all the training materials will require hefty backing. Additionally that if you download a model file that contains enough of the source material to be considered infringing (even without using the LLM, assume you can extract the contents directly out of the weights) then it might as well be a .zip with a PDF in it, the mode…

This technology is a really bad way of storing, reproducing and transmitting the books themselves. It's probabilistic and lossy. It may be possible to reproduce some paragraphs, but no reasonable person would expect to read The Da Vinci Code by prompting the LLM. Surely the marketed use cases and the observed real use by users has to make it clear that the intended and vastly overwhelming use of an LLM is transformat…

Not The DaVinci Code, but I recently tried reading "OCaml Programming: Correct + Efficient + Beautiful" through Gemini. The book is open, so I rightly assumed it was "in there". I read by saying "Give me the first paragraph of Chapter 6" and then something like "Next 3 paragraphs". If I had a question, I was able to ask it and get some more info and have something like a dialog.

As far as I could tell, the book didn't match what's posted online today. The text was somewhat consistent on a topic, yet poorly written and made references to sections that I don't think existed. No amount of prompting could locate them. I'm not convinced the material presented to me was actually the book, although it seemed consistent with the topic of the chapter.

I tried to ascertain when the book had been scraped, yet couldn't find a match in Archive.org or in the book's git repo.

Eventually I gave up and just continued reading the PDF.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#146

Earlier quoted context omitted.

> get fair compensation for the amount of work This is a bit distorted. This is a better summary: The primary purpose of copyright is to induce and reward authors to create new works and to make those works available to the public to enjoy. The ultimate purpose is to foster the creation of new works that the public can read and written culture can thrive. The means to achieve this is by ensuring that the authors of s…

As you point out, people make rules ("laws") which benefit them . I care about fairness and justice though, even if I am a minority. Fundamentally, fair compensation is based on the amount of work put in (obviously taking skill/competence into account but the differences between people in most disciplines probably don't span a single order of magnitude, let alone several). The ultimate goal should be to prevent peopl…

The "you wouldn't download a car" argument made with a straight face. Remarkable.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#147

Earlier quoted context omitted.

Yep, broadly capable open models are on track for annihilation. The cost of legally obtaining all the training materials will require hefty backing. This will have the effect of empowering countries (and other entities) that don't respect copyright law, of course. The copyright cartel cannot be allowed to yank the handbrake on AI. If they insist on a fight, they must lose.

For that matter, how dare the government fine me for dumping waste in the river, and stop me from employing minors? Don't they know it will ruin the economy?

Copyright is something we invented from thin air, and relatively recently at that. Meanwhile, refraining from fouling their own nests is something that most animals have accomplished instinctually for millions of years.

So, not really comparable.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#148

Earlier quoted context omitted.

Yep, broadly capable open models are on track for annihilation. The cost of legally obtaining all the training materials will require hefty backing. Additionally that if you download a model file that contains enough of the source material to be considered infringing (even without using the LLM, assume you can extract the contents directly out of the weights) then it might as well be a .zip with a PDF in it, the mode…

Yep, broadly capable open models are on track for annihilation. The cost of legally obtaining all the training materials will require hefty backing. This will have the effect of empowering countries (and other entities) that don't respect copyright law, of course. The copyright cartel cannot be allowed to yank the handbrake on AI. If they insist on a fight, they must lose.

[flagged]

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#149

Earlier quoted context omitted.

Yep, broadly capable open models are on track for annihilation. The cost of legally obtaining all the training materials will require hefty backing. This will have the effect of empowering countries (and other entities) that don't respect copyright law, of course. The copyright cartel cannot be allowed to yank the handbrake on AI. If they insist on a fight, they must lose.

[flagged]

We spun "intellectual property" law from whole cloth. We'll need to reweave it now. Deal with it.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#150

Earlier quoted context omitted.

> get fair compensation for the amount of work This is a bit distorted. This is a better summary: The primary purpose of copyright is to induce and reward authors to create new works and to make those works available to the public to enjoy. The ultimate purpose is to foster the creation of new works that the public can read and written culture can thrive. The means to achieve this is by ensuring that the authors of s…

As you point out, people make rules ("laws") which benefit them . I care about fairness and justice though, even if I am a minority. Fundamentally, fair compensation is based on the amount of work put in (obviously taking skill/competence into account but the differences between people in most disciplines probably don't span a single order of magnitude, let alone several). The ultimate goal should be to prevent peopl…

I understand your sense of justice in cheering on David against Goliath. But the equation is not so clear. The common person is sometimes on this side, sometimes on that side. Copyright can also be weaponized by megacorps against normal people (copying Disney movie DVDs) and LLMs can also be in the hands of the decentralized public (llama ecosystem).

The house thing is a bit offtopic because to be considered for copyright, only its artistic, architectural expression matters. If you want to protect the ingenuity in the technical ways of how it's constructed, that's a patent law thing. It also muddies the water by bringing in aspects of the privacy of one's home by making us imagine paparazzi style photoshoots and sneaky X rays.

The thing is, houses can't be copied like bits and bytes. I would copy a car if I could. If you could copy a loaf of bread for free, it would be a moral imperative to do so, whatever the baker might think about it.

> fair compensation is based on the amount of work put in

This is the labor theory of value, but it has many known problems. For example that the amount of work put in can be disconnected from the amount of value it provides to someone. Pricing via supply/demand market forces have produced much better outcomes across the globe than any other type of allocation. Of course moderated by taxes and so on.

But overall the question is whether LLMs create value for the public. Does it foster prosperity of society? If yes, laws should be such that LLMs can digest more books rather than less. If LLMs are good, they should not be restricted to be trained on copyright-expired writings.

Post reply on HN