Live data from Hacker News

A federal judge sides with Anthropic in lawsuit over training AI on books

techcrunch.com

111–120 of 222 posts

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#111
post #19

Earlier quoted context omitted.

What if I overfit my LLM so it spits out copyrighted work with special prompting? Where to draw the line in training?

I mean the human brain can memorize things as well and it’s not illegal. It’s only illegal if said memorized thing is distributed.

Humans don't scale. LLMs do.

Even if LLMs were actual human-level AI (they are not - by far), a small bunch of rich people could use them to make enormous amounts of money without putting in the enormous amounts of work humans would have to.

All the while "training" (= precomputing transformations which among other things make plagiarism detection difficult) on work which took enormous amounts of human labor without compensating those workers.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#112
post #62

Earlier quoted context omitted.

They will infringe copyright as soon as they are sufficiently similar to the original. You can’t shoot a non-verbatim but clearly recognizable beat-by-beat remake of Star Wars, call it Galaxy Conflict, and get away with monetizing it.

Correct. You have to call it "Starcrash" ( https://www.imdb.com/title/tt0079946/?ref_=ls_t_8 ). Then it's legal.

Interesting artifact, but the very first/top IMDB user review convincingly contradicts that this is a Star Wars remake. ;)

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#113

Earlier quoted context omitted.

The analogy to training is not writing a play based on the work. It's more like reading (experiencing) the work and forming memories in your brain, which you can access later. I'm allowed to hear a copyrighted tune, and even whistle it later for my own enjoyment, but I can't perform it for others without license.

This is nonsense, in my opinion. You aren't "hearing" anything. You are literally creating a work, in this case, the model, derived from another work. People need to stop anthropomorphizing neural networks. It's a software and a software is a tool and a tool is used by a human.

Humans are also created/derived from other works, trained, and used as a tool by humans.

It's interesting how polarizing the comparison of human and machine learning can be.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#114

One aspect of this ruling [1] that I find concerning: on pages 7 and 11-12, it concedes that the LLM does substantially "memorize" copyrighted works, but rules that this doesn't violate the author's copyright because Anthropic has server-side filtering to avoid reproducing memorized text. (Alsup compares this to Google Books, which has server-side searchable full-text copies of copyrighted books, but only allows user…

Copyright was codified in an age where plagiarism was time consuming. Even replacing words with synonyms on a mass scale was technically infeasible. The goal of copyright is to make sure people can get fair compensation for the amount of work they put in. LLMs automate plagiarism on a previously unfathomable scale. If humans spend a trillion hours writing books, articles, blog posts and code, then somebody (a small g…

> get fair compensation for the amount of work

This is a bit distorted. This is a better summary: The primary purpose of copyright is to induce and reward authors to create new works and to make those works available to the public to enjoy.

The ultimate purpose is to foster the creation of new works that the public can read and written culture can thrive. The means to achieve this is by ensuring that the authors of said works can get financial incentives for writing.

The two are not in opposition but it's good to be clear about it. The main beneficiary is intended to be the public, not the writers' guild.

Therefore when some new factor enters the picture such as LLMs, we have to step back and see how the intent to benefit the reading public can be pursued in the new situation. It certainly has to take into account who and how will produce new written works, but it is not the main target, but can be an instrumental subgoal.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#115

Earlier quoted context omitted.

https://news.ycombinator.com/item?id=44369227 If the US makes it illegal to train LLMs on copyrighted data, the US will find a solution and not just give up and wait half a decade to see what China does in the meantime.

What solution is there?

Zillow have the MLSs network that provide them lists, a similar solution could apply if courts agree that library copies apply for this - Anthropic could sign agreements with large libraries and "check out"/"freeze" copies for a minimally-agreed-upon duration and query across all to see which has a copy of each book they need. Spotify and Apple Music sign deals en masse with labels, the same could apply here with book publishers, labels for lyrics, museums for art, etc. Or whatever other creative solution that people who will need to find, will find. Right now they took the laziest path, because it worked. They will find the next-laziest path that works.

And the easiest option: Legislation change. If it's completely decided that the current law blocks LLMs from working in the US, the industry will lobby to amend the copyright law (which is not immutable) to add a carveout for it.

You're assuming that people will just give up. People never gave up, why would they now?

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#116

I'm surprised we never discuss a previous case of how governments handled a valuable new technology that challenged creative's ability to monetise their work: Cassette Tapes and Private Copying Levy. https://en.wikipedia.org/wiki/Private_copying_levy Governments didn't ban tapes but taxed them and fed the proceeds back into the royalty system. An equivalent for books might be an LLM tax funding a negative tax rate fo…

Surely this would require the observation that the public is actually using LLMs as a substitute for purchasing the book, ie they sit down and type "Generate me the first/second/third chapter of The Da Vinci Code" and then read if from there. Because it was easy to observe in the cassette tape era that people copied the store bought music and films and shared it among each other. I doubt that this is or will be a serious use case of LLMs.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#117
post #73

Will be interesting to see how this affects Anthropic's ongoing lawsuit with Reddit, or all the different media publishing ones flying around. Is it okay to train on books but not online posts and articles? Why the distinction between the two?

The distinction will be whether those online posts were obtained legally, analogous to whether the books in this case were pirated.

It’s not as simple as it sounds, since I’m sure scraping is against Reddit’s terms and conditions, but if those posts are made publicly available without the scraper actually agreeing to anything, is that a valid breach of contract?

Will be interesting to see how that plays out.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#118

> “We will have a trial on the pirated copies used to create Anthropic’s central library and the resulting damages,” Judge Alsup wrote in the decision. “That Anthropic later bought a copy of a book it earlier stole off the internet will not absolve it of liability for theft but it may affect the extent of statutory damages.” I'm not sure why this alone is considered a separate issue from training the AI with books. B…

The ruling suggests that "pirating a book that could have been bought at a bookstore" for the sake of "writing a book review" "is inherently, irredeemably infringing".

Which suggests that, at least in the judge's opinion, 'fair use rights' do exist in a sense, but it's about when you read the book, not when you publish.

But that's not settled precedent. Meta is currently arguing the opposite in Kadrey v. Meta: they're claiming that they can get away with torrenting training material as long as they only leech (download) and don't seed (upload), because, although the act of downloading (copying) is generally infringement under a Ninth Circuit precedent, they were making a fair use.

As for EULAs, that might be true for e-books, but publishers can't really do anything about Anthropic's new strategy of scanning physical books, because physical books generally don't come with shrinkwrap license agreements. Perhaps publishers could start adding them, but I think that would sit poorly with the public and the courts.

(That's assuming the ruling isn't overturned on appeal, which it easily might be.)

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#119

I'm surprised we never discuss a previous case of how governments handled a valuable new technology that challenged creative's ability to monetise their work: Cassette Tapes and Private Copying Levy. https://en.wikipedia.org/wiki/Private_copying_levy Governments didn't ban tapes but taxed them and fed the proceeds back into the royalty system. An equivalent for books might be an LLM tax funding a negative tax rate fo…

That's a very different use case IMO. An LLM isn't generating a replica of a book for the users. At most we've seen people able to reproduce exact portions of stuff, but only with lots of prior knowledge of the material by the human in the loop and plenty of manual effort (aka not a direct commercial threat). And that was before more LLMs put effort into stopping that sort of hacking.

The last thing the world needs is more nonsensical copyright law and hand wavy regulation funded by entrenched interests.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#120
Humans read books. AI/LLMs do not read. I think there's an inherent difference here. If the LLM is making a copy of the entire book in it's memory, is that copyright infringement? I don't know the answer to that, but it feels like Alsup is considering this fair use argument in the context of a human, but it's nothing like a human and needs to be treated differently.
Post reply on HN