Live data from Hacker News

A federal judge sides with Anthropic in lawsuit over training AI on books

techcrunch.com

71–80 of 222 posts

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#71
post #3

The HN crowd dislikes brick-and-mortar landlords but often sides with charging rent for certain bits. Which side will prevail? Interesting excerpt: > “We will have a trial on the pirated copies used to create Anthropic’s central library and the resulting damages,” Judge Alsup wrote in the decision. “That Anthropic later bought a copy of a book it earlier stole off the internet will not absolve it of liability for the…

I think the reason it's okay to charge rent for certain bits is that the space of bitstrings is so large.

Choosing someone's bitstrings is like choosing to harvest someone's fields in a world where there's infinite space of fertile fields. You picked his, instead of finding a space in the infinite expanse to farm on your own.

If you start writing something you'll never generate a copyrighted work at random. When the work isn't available nothing is taken away from you even if you were strictly forbidden from reproducing the work.

Choosing someone's particular bitstring is only done because there's someone who has expended effort in preparing it.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#72
post #64
post #38

Earlier quoted context omitted.

> But those training the LLMs are still using the works, and not just to discuss them, which I think is the point of fair use doctrine. Worse, they’re using it for massive commercial gain, without paying a dime upstream to the supply chain that made it possible. If there is any purpose of copyright at all, it’s to prevent making money from someone’s else’s intellectual work. The entire thing is based on economic prag…

> Worse, they’re using it for massive commercial gain, without paying a dime upstream to the supply chain that made it possible. If there is any purpose of copyright at all, it’s to prevent making money from someone’s else’s intellectual work. This makes no sense. If I buy and read a book on software engineering, and then use that knowledge to start a career, do I owe the author a percentage of my lifetime earnings?…

If you pirate a book on software engineering and then use that knowledge to start a career, do you owe the author the royalties they would be paid had you bought the book?

If the career you start isn't software engineering directly but instead re-teaching the information you learned from that book to millions of paying students, is the regular royalty payment for the book still fair?

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#73
Will be interesting to see how this affects Anthropic's ongoing lawsuit with Reddit, or all the different media publishing ones flying around. Is it okay to train on books but not online posts and articles? Why the distinction between the two?

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#74
post #2

Broadly summarizing. This is OK and fair use: Training LLMs on copyrighted work, since it's transformative. This is not OK and not fair use: pirating data, or creating a big repository of pirated data that isn't necessarily for AI training. Overall seems like a pretty reasonable ruling?

It’s similar to the Google Books ruling, which Google lost. Anthropic also lost. TechCrunch and others are very aspirational here.

Do you mean Authors Guild, Inc. v. Google, Inc.? Google won that case:

https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....

Maybe there's another big Google Books lawsuit that Google ultimately lost, but I don't know which one you mean in that case.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#75

One aspect of this ruling [1] that I find concerning: on pages 7 and 11-12, it concedes that the LLM does substantially "memorize" copyrighted works, but rules that this doesn't violate the author's copyright because Anthropic has server-side filtering to avoid reproducing memorized text. (Alsup compares this to Google Books, which has server-side searchable full-text copies of copyrighted books, but only allows user…

Wouldn’t the issue be executing the models to third parties without filters? No idea if this is right but the same it would apply to Anthropic that they couldn’t run the model without the filter system having a chicken an egg problem. Can’t develop the filter without looking into the model.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#76
post #53

Earlier quoted context omitted.

It might be different if you are a commercial product which couldn’t have been created without incorporating the contents of all those books. Humans, animals, hardware and software are treated differently by law because they have different constraints and capabilities.

But a commercial product is reaching parity with human capability. Let's be real, Humans have special treatment (more special than animals as we can eat and slaughter animals but not other humans) because WE created the law to serve humans. So in terms of being fair across the board LLMs are no different. But there's no harm in giving ourselves special treatment.

Generative AIs are very different from humans because they can be copied losslessly and scaled tremendously, and also have no individual liability, nor awareness of how similar their output is to something in their training material. They are very different in constraints and capabilities from humans in all sorts of ways. For one, a human will likely never reproduce a book they read without being aware that that’s what they are doing.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#77
post #19

Earlier quoted context omitted.

What if I overfit my LLM so it spits out copyrighted work with special prompting? Where to draw the line in training?

I mean the human brain can memorize things as well and it’s not illegal. It’s only illegal if said memorized thing is distributed.

Because humans have rights

AI models do not.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#78

Earlier quoted context omitted.

Definitely seems reasonable to say "you can train on this data but you have to have a legal copy" Personally I like to frame most AI problems by substituting a human (or humans) for the AI. Works pretty well most of the time. In this case if you hired a bunch of artists/writers that somehow had never seen a Disney movie and to train them to make crappy Disney clones you made them watch all the movies it certainly wou…

That's a part of the issue. I'm not sure if this has happened in visual arts, but there is in fact precedent against trying to hire a sound a like over the one you want to sound like. You can't be in talks with Scarlet Johannsen, reject her, and then hire a sound a like and say "talk like Scarlet". It's pretty clear at that point what you want but you didn't want to pay talent for it. I see elements of that here. Buy…

> but there is in fact precedent against trying to hire a sound a like over the one you want to sound like. You can't be in talks with Scarlet Johannsen, reject her, and then hire a sound a like and say "talk like Scarlet". It's pretty clear at that point what you want but you didn't want to pay talent for it.

You're referencing Midler v Ford Motor Co in the 9th circuit. This case largely applies to California, not the whole nation. Even then, it would take one Supreme Court case to overturn it.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#79
post #10

Earlier quoted context omitted.

But those training the LLMs are still using the works, and not just to discuss them, which I think is the point of fair use doctrine. I guess I fail to see how it's any different from me using it in some other way? If I wanted to write a play very loosely inspired by Blood Meridian, it might be transformative, but that doesn't justify me pirating the book. I tend to think copyright should be extremely limited compare…

The analogy to training is not writing a play based on the work. It's more like reading (experiencing) the work and forming memories in your brain, which you can access later. I'm allowed to hear a copyrighted tune, and even whistle it later for my own enjoyment, but I can't perform it for others without license.

This is nonsense, in my opinion. You aren't "hearing" anything. You are literally creating a work, in this case, the model, derived from another work.

People need to stop anthropomorphizing neural networks. It's a software and a software is a tool and a tool is used by a human.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#80
post #2

Broadly summarizing. This is OK and fair use: Training LLMs on copyrighted work, since it's transformative. This is not OK and not fair use: pirating data, or creating a big repository of pirated data that isn't necessarily for AI training. Overall seems like a pretty reasonable ruling?

If a publisher adds a "no AI training" clause to their contracts, does this ruling render it invalid?

You don't need a license for most of what people do with traditional, physical copyrighted copies of works: read them, play a DVD at home, etc. Those things are outside the scope of copyright. But you do need a license to make copies, and ebooks generally come with licensing agreements, again because to read an ebook, you must first make a brand new copy of it. Anyway as a result physical books just don't have "licenses" to begin with and if they tried they'd be unenforceable, since you don't need to "agree" to any "terms" to read a book.
Post reply on HN