Live data from Hacker News

A federal judge sides with Anthropic in lawsuit over training AI on books

techcrunch.com

161–170 of 222 posts

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#161

Earlier quoted context omitted.

It's not a copyright violation to summarize (in different words).

A quick Google search will reveal that this not the case. Summaries of books or movies have no particular legal protection and the authors of those summaries may be sued by the owners of that content. https://1minutebook.com/are-book-summaries-legal/ Fair use is a defense often cited in those cases but it's just that: a defense. Cliff Notes are often cited here but they actually license the content in many cases.

I mean, have you actually read the text at the link you provided? Or just remembered something, googled quickly and sent a random hit without reading it? The quotes under "What do lawyers say? Listen to what a several Intellectual Property Lawyers are saying on “Are book summaries legal?”:" certainly seem to be closer to what I was claiming.

> If you want to write a summary of any novel, without quoting from it, you are free to do it

> Copyright does not protect ideas, only a particular expression of those ideas

> You would likely get in trouble only if your summary contained long excerpts directly from the book

> As long as you do not quote directly from the book, or copy any of the content, then writing a unique summary is not illegal. You can mention the title, you can even quote sentences from the book as long as they are cited, you just can’t reproduce chunks of the content

etc

(I'm also not sure whether this article is just blogspam or itself AI generated)

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#162

Earlier quoted context omitted.

Cliff Notes are fair use. Would you argue otherwise? Wikipedia also has plot summaries without infringement.

In your parent comment, you argued what people would do in practice. Now you have shifted to talking about what is legal or not to do. I'm not a legal scholar, so I'm not qualified or interested in arguing about whether Cliff Notes is fair use. But I do care about how people behave, and I'm pretty sure that Cliff Notes and LLMs lead to fewer books being purchased, which makes it harder for writers to do what they do.…

It surely matters whether people actually use the thing for copyright violations or not. Summaries are not even copyright violations so that's irrelevant. Long verbatim copies would be, but one would have to demonstrate that this use case is significant, convenient enough to provide a viable alternative to otherwise obtaining the particular text chunk etc.

----

> But for authors of newer technical material, yes, I think LLMs will make it harder for those people to be able to afford to spend the time thinking, writing, and sharing their expertise.

Alright, you're now arguing for some new regulations though, since this is not a matter for copyright.

In that context, I observe that many academics already put their technical books online for free. Machine learning, computer vision, robotics etc. I doubt it's a hugely lucrative thing in the first place.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#163

Earlier quoted context omitted.

Yep, broadly capable open models are on track for annihilation. The cost of legally obtaining all the training materials will require hefty backing. Additionally that if you download a model file that contains enough of the source material to be considered infringing (even without using the LLM, assume you can extract the contents directly out of the weights) then it might as well be a .zip with a PDF in it, the mode…

> broadly capable open models are on track for annihilation I'm not so sure about this one. In particular, presuming that it is found that models which can produce infringing material are themselves infringing material, the ability to distill models from older models seems to suggest that the older models can actually produce the new, infringing model. It seems like that should mean that all output from the older mod…

In theory, couldn't you distill a non-infringing model from an infringing one? Just prompt it for continuations and give it a whack every time the output matches something in your dataset of copyrighted works.

You'd need the copyrighted works to compare to, of course, though if you have the permissible training data (as Anthropic apparently does) it should be doable.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#165

Earlier quoted context omitted.

It's not a copyright violation to summarize (in different words).

The impact of machinary forces re-evaluation of any concepts defined in terms of human capability because scale/automation changes their nature. Just as duplicating a fragment can be legal, duplicating any fragment on demand is not. Rephrasing a passage might be legal, but rephrasing any passage on demand might not.

That's reasonable. This would require broader and deeper thought and discussion apart from the strict legal debate. As in, what is the public interest here? What kinds of rules would bring social good? Etc. What should the law facilitate and what should it limit to achieve that? The problem is, that we really don't know how things will play out, we have no long-term experience with these things yet. So it's all very speculative.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#166
post #104

One aspect of this ruling [1] that I find concerning: on pages 7 and 11-12, it concedes that the LLM does substantially "memorize" copyrighted works, but rules that this doesn't violate the author's copyright because Anthropic has server-side filtering to avoid reproducing memorized text. (Alsup compares this to Google Books, which has server-side searchable full-text copies of copyrighted books, but only allows user…

Yes and no. In this case, the plaintiffs alleged that Anthropic's LLMs had memorized the works so completely that "if each completed LLM had been asked to recite works it had trained upon, it could have done so", "almost verbatim". The judge assumed for the sake of argument that the allegation was true, and ruled that the conduct was fair use anyway due to the existence of an effective filter. Therefore there was no…

Wouldn't a model that can recite training data verbatim be larger than necessary? Exact text isn't coming from nowhere, no matter how efficiently the bits are encoded, and the same effectiveness should be achievable by compressing those portions of the model.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#167

Earlier quoted context omitted.

[flagged]

We spun "intellectual property" law from whole cloth. We'll need to reweave it now. Deal with it.

Rewriting does not mean destroying, as the cannibalization of news reporting by social media should have taught us.

It's entirely possible for something to be suboptimal in the specific (I would like this thing for free), but optimal on the whole (society benefits from this thing not being free).

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#168

Earlier quoted context omitted.

As you point out, people make rules ("laws") which benefit them . I care about fairness and justice though, even if I am a minority. Fundamentally, fair compensation is based on the amount of work put in (obviously taking skill/competence into account but the differences between people in most disciplines probably don't span a single order of magnitude, let alone several). The ultimate goal should be to prevent peopl…

I understand your sense of justice in cheering on David against Goliath. But the equation is not so clear. The common person is sometimes on this side, sometimes on that side. Copyright can also be weaponized by megacorps against normal people (copying Disney movie DVDs) and LLMs can also be in the hands of the decentralized public (llama ecosystem). The house thing is a bit offtopic because to be considered for copy…

The "fairness" argument is weaker than the "sustainable creation" one.

If LLMs could create quality literature, or social media create in-depth reporting, then I'd have no problem with the tide of technological progress flowing.

Unfortunately, recent history has shown that it's trivial for the market to cannibalize the financial model of creators without replacing it.

And as a result, society gets {no more that thing} + {watered down, shitty version}.

Which isn't great.

So I'd love to hear an argument from the 'fuck copyright, let's go AI' crowd (not the position you seem to be espousing) on what year +10 of rampant AI ingestion of copyrighted works looks like...

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#169

Earlier quoted context omitted.

It's different but not in ways that make such interventions irrelevant e.g. why would we only care about lost sales? If copyright has been violated as a necessary means to generate new value, haven't the content creators earned this value? Such imperfect measures offer a compromise between "big tech can steal everything" and "LLMs trained on unpurchased books are illegal". It's not just books but any tragedy-of-the-c…

> It's different but not in ways that make such interventions irrelevant e.g. why would we only care about lost sales? If copyright has been violated as a necessary means to generate new value, haven't the content creators earned this value? Indeed the company should purchase the books. If they obtain copies in a process that violates copyright, then that's indeed a violation of copyright. The current decision does n…

Anthropic apparently did it both ways. After realizing that pirating mass quantities of books for training wasn't a great legal look, it hired someone previously responsible for Google Books, who in turn contacted publishers about mass licensing their content for training use.

However, that option was ultimately not pursued as instead...

>> Anthropic spent many millions of dollars to purchase millions of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals. Each print book resulted in a PDF copy containing images of the scanned pages with machine-readable text (including front and back cover scans for softcover books). Anthropic created its own catalog of bibliographic metadata for the books it was acquiring. It acquired copies of millions of books, including of all works at issue for all Authors.

(from the ruling)

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#170

Earlier quoted context omitted.

Transformed summaries are generally fair use already (or perhaps not even an issue of copyright). You can read plot summaries of novels and movies on Wikipedia, same with technical topics. The ideas are not protected by copyright, the artistic expression is. Certain technical ideas can be protected via patents. But even then, not the description of idea, but putting it into practice. Ideas that you're not supposed to…

Are you sure? Or are owners deciding not to sue because they are seeing some benefit? I believe copyright is always case-by-case. No one sues over plot summaries because they likely help sales. Summarize books or news articles with an LLM and you end up with the lawsuits we see today.

The specific difference is summarizing automatically, at scale, which is a novel technological possibility.

The previous balance of rights was created when summarizing took human time and proceeded at human pace.

Now, that's different and a new balance needs to be struck.

Post reply on HN