Live data from Hacker News

A federal judge sides with Anthropic in lawsuit over training AI on books

techcrunch.com

21–30 of 222 posts

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#21
post #2

Broadly summarizing. This is OK and fair use: Training LLMs on copyrighted work, since it's transformative. This is not OK and not fair use: pirating data, or creating a big repository of pirated data that isn't necessarily for AI training. Overall seems like a pretty reasonable ruling?

Definitely seems reasonable to say "you can train on this data but you have to have a legal copy" Personally I like to frame most AI problems by substituting a human (or humans) for the AI. Works pretty well most of the time. In this case if you hired a bunch of artists/writers that somehow had never seen a Disney movie and to train them to make crappy Disney clones you made them watch all the movies it certainly wou…

That's a part of the issue. I'm not sure if this has happened in visual arts, but there is in fact precedent against trying to hire a sound a like over the one you want to sound like. You can't be in talks with Scarlet Johannsen, reject her, and then hire a sound a like and say "talk like Scarlet". It's pretty clear at that point what you want but you didn't want to pay talent for it.

I see elements of that here. Buying copyrighted works not to be exposed and be inspired, nor to utilize the aithor's talents, but to fuel a commercialization of sound-a-likes.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#22
post #2

Broadly summarizing. This is OK and fair use: Training LLMs on copyrighted work, since it's transformative. This is not OK and not fair use: pirating data, or creating a big repository of pirated data that isn't necessarily for AI training. Overall seems like a pretty reasonable ruling?

If a publisher adds a "no AI training" clause to their contracts, does this ruling render it invalid?

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#24
post #2

Broadly summarizing. This is OK and fair use: Training LLMs on copyrighted work, since it's transformative. This is not OK and not fair use: pirating data, or creating a big repository of pirated data that isn't necessarily for AI training. Overall seems like a pretty reasonable ruling?

If a publisher adds a "no AI training" clause to their contracts, does this ruling render it invalid?

what contract? with who?

Meta at least just downloaded ENGLISH_LANGUAGUE_BOOKS_ALL_MEGATORRENT.torrent and trained on that.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#25

Earlier quoted context omitted.

Definitely seems reasonable to say "you can train on this data but you have to have a legal copy" Personally I like to frame most AI problems by substituting a human (or humans) for the AI. Works pretty well most of the time. In this case if you hired a bunch of artists/writers that somehow had never seen a Disney movie and to train them to make crappy Disney clones you made them watch all the movies it certainly wou…

That's a part of the issue. I'm not sure if this has happened in visual arts, but there is in fact precedent against trying to hire a sound a like over the one you want to sound like. You can't be in talks with Scarlet Johannsen, reject her, and then hire a sound a like and say "talk like Scarlet". It's pretty clear at that point what you want but you didn't want to pay talent for it. I see elements of that here. Buy…

> You can't be in talks with Scarlet Johannsen, reject her, and then hire a sound a like and say "talk like Scarlet"

Keep in mind, the Authors in the lawsuit are not claiming the _output_ is copyright infringement so Alsup isn't deciding that.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#26
post #2

Broadly summarizing. This is OK and fair use: Training LLMs on copyrighted work, since it's transformative. This is not OK and not fair use: pirating data, or creating a big repository of pirated data that isn't necessarily for AI training. Overall seems like a pretty reasonable ruling?

Definitely seems reasonable to say "you can train on this data but you have to have a legal copy" Personally I like to frame most AI problems by substituting a human (or humans) for the AI. Works pretty well most of the time. In this case if you hired a bunch of artists/writers that somehow had never seen a Disney movie and to train them to make crappy Disney clones you made them watch all the movies it certainly wou…

> Definitely seems reasonable to say "you can train on this data but you have to have a legal copy"

How many copies? They're not serving a single client.

Libraries need to have multiple e-book licenses, after all.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#27
post #2

Broadly summarizing. This is OK and fair use: Training LLMs on copyrighted work, since it's transformative. This is not OK and not fair use: pirating data, or creating a big repository of pirated data that isn't necessarily for AI training. Overall seems like a pretty reasonable ruling?

BRB, I'm going to download all the TV shows and movies to train my vision model. Just to be sure it's working properly, I have to watch some for debugging purposes.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#28
post #7

Earlier quoted context omitted.

For textual purposes it seems fairly transformative. If you train a LLM on harry potter and ask it to generate a story that isn't harry potter then it's not a replacement. However, if you train a model on stock imagery and use it to generate stock imagery then I think you'll run into an issue from the Warhol case.

The nature of how they store data makes it not okay in my books. You massage the data enough and you can generate something that seems infringement worthy.

For closed models the storage problem isn't really a problem, they can be judged by what they produce not how they store it as you don't have access to the actual data. That said, open weight LLMs are probably screwed, if enough of the work remains in the weights such that they can be extracted (even if it's without even talking to the LLM) then the weight file itself represents a copy of the work that's being distributed. So enjoy these competent run-at-home models while you can, they're on track for extinction.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#29
post #2

Broadly summarizing. This is OK and fair use: Training LLMs on copyrighted work, since it's transformative. This is not OK and not fair use: pirating data, or creating a big repository of pirated data that isn't necessarily for AI training. Overall seems like a pretty reasonable ruling?

If a publisher adds a "no AI training" clause to their contracts, does this ruling render it invalid?

Fair use overrides licensing

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#30
post #14
post #10

Earlier quoted context omitted.

But those training the LLMs are still using the works, and not just to discuss them, which I think is the point of fair use doctrine. I guess I fail to see how it's any different from me using it in some other way? If I wanted to write a play very loosely inspired by Blood Meridian, it might be transformative, but that doesn't justify me pirating the book. I tend to think copyright should be extremely limited compare…

>If I wanted to write a play very loosely inspired by Blood Meridian, it might be transformative, but that doesn't justify me pirating the book. I think that's the conclusion of the judge. If Anthropic were to buy the books and train on them, without extra permission from the authors, it would be fair use, much like if you were to be inspired by it (though in that case, it may not even count as a derivative work at a…

which AFAIN IANAL, copyright and exhaustive rights are completely different. Under copyright, once a book is purchased: that's it. Reselling the same, or transformed (re: highlighted) worked 'used' is 100% legal, as is consuming it at your discretion (in your mind {a billion times}, a fire, or (yes even) what amounts to a fancy calculator).

(that's all to say copyright is dated and needs an overhaul)

But that's taking a viewpoint of 'training a personal AI in your home', which isn't something that actually happens... The issue has never been the training data itself. Training an AI and 'looking at data and optimizing a (human understanding/AI understanding) function over it' are categorically the same, even if mechanically/biologically they are very different.

Post reply on HN