Live data from Hacker News

A federal judge sides with Anthropic in lawsuit over training AI on books

techcrunch.com

11–20 of 222 posts

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#11
post #4
post #2

Broadly summarizing. This is OK and fair use: Training LLMs on copyrighted work, since it's transformative. This is not OK and not fair use: pirating data, or creating a big repository of pirated data that isn't necessarily for AI training. Overall seems like a pretty reasonable ruling?

Depends whether you actually agree its transformative

What's the steelman case that is transformative? Because prima-facie, it seems to only output original output - "intelligent" output.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#12
post #7
post #4

Earlier quoted context omitted.

Depends whether you actually agree its transformative

For textual purposes it seems fairly transformative. If you train a LLM on harry potter and ask it to generate a story that isn't harry potter then it's not a replacement. However, if you train a model on stock imagery and use it to generate stock imagery then I think you'll run into an issue from the Warhol case.

The nature of how they store data makes it not okay in my books. You massage the data enough and you can generate something that seems infringement worthy.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#13
post #2

Broadly summarizing. This is OK and fair use: Training LLMs on copyrighted work, since it's transformative. This is not OK and not fair use: pirating data, or creating a big repository of pirated data that isn't necessarily for AI training. Overall seems like a pretty reasonable ruling?

Definitely seems reasonable to say "you can train on this data but you have to have a legal copy"

Personally I like to frame most AI problems by substituting a human (or humans) for the AI. Works pretty well most of the time.

In this case if you hired a bunch of artists/writers that somehow had never seen a Disney movie and to train them to make crappy Disney clones you made them watch all the movies it certainly would be legal to do so but only if they had legit copies in the training room. Pirating the movies would be illegal.

Though the downside is it does create a training moat. If you want to create the super-brain AI that's conversant on the corpus of copyrighted human literature you're going to need a training library worth millions

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#14
post #10
post #2

Broadly summarizing. This is OK and fair use: Training LLMs on copyrighted work, since it's transformative. This is not OK and not fair use: pirating data, or creating a big repository of pirated data that isn't necessarily for AI training. Overall seems like a pretty reasonable ruling?

But those training the LLMs are still using the works, and not just to discuss them, which I think is the point of fair use doctrine. I guess I fail to see how it's any different from me using it in some other way? If I wanted to write a play very loosely inspired by Blood Meridian, it might be transformative, but that doesn't justify me pirating the book. I tend to think copyright should be extremely limited compare…

>If I wanted to write a play very loosely inspired by Blood Meridian, it might be transformative, but that doesn't justify me pirating the book.

I think that's the conclusion of the judge. If Anthropic were to buy the books and train on them, without extra permission from the authors, it would be fair use, much like if you were to be inspired by it (though in that case, it may not even count as a derivative work at all, if the relationship is sufficiently loose). But that doesn't mean they are free to pirate it either, so they are likely to be liable for that (exactly how that interpretation works with copyright law I'm not entirely sure: I know in some places that downloading stuff is less of a problem than distributing it to others because the latter is the main thing that copyright is concerned with. And AFAIK most companies doing large model training are maintaining that fair use also extends to them gathering the data in the first place).

(Fair use isn't just for discussion. It covers a broad range of potential use cases, and they're not enumerated precisely in copyright law AFAIK, there's a complicated range of case law that forms the guidelines for it)

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#16
post #7
post #4

Earlier quoted context omitted.

Depends whether you actually agree its transformative

For textual purposes it seems fairly transformative. If you train a LLM on harry potter and ask it to generate a story that isn't harry potter then it's not a replacement. However, if you train a model on stock imagery and use it to generate stock imagery then I think you'll run into an issue from the Warhol case.

Wasn't that just over an arrangement of someone else's photographs?

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#17
post #6

Earlier quoted context omitted.

Anthropic won't submit a spreadsheet of all the books and whether they were purchases or not. So trivially, not every book stolen is shown to be later purchased. As just a matter of society, I don't think you want people say stealing a car and then coming back a month later with the money.

While no one wants anyone to steal a car, almost no one would mind freely cloning a car. The trouble truly is that 3d-printing hasn't gotten that good yet.

If 3d printing was that good, stealing a car would be moot because production costs would come way down and only need to cover cost/procurement of materials and paying back the black box.

Regardless, I don't think the car is an apt metaphor here. Cars are an important utility and gatekeeping cars arguably holds society back., art is creative expression, and no one is going hungry because they didn't have $10 for the newest book.

We also have libraries already for this reason, so why not expand on that instead of relinquishing sharing of knowledge to a private corporation?

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#18
post #2

Broadly summarizing. This is OK and fair use: Training LLMs on copyrighted work, since it's transformative. This is not OK and not fair use: pirating data, or creating a big repository of pirated data that isn't necessarily for AI training. Overall seems like a pretty reasonable ruling?

It’s similar to the Google Books ruling, which Google lost. Anthropic also lost. TechCrunch and others are very aspirational here.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#19
post #2

Broadly summarizing. This is OK and fair use: Training LLMs on copyrighted work, since it's transformative. This is not OK and not fair use: pirating data, or creating a big repository of pirated data that isn't necessarily for AI training. Overall seems like a pretty reasonable ruling?

What if I overfit my LLM so it spits out copyrighted work with special prompting? Where to draw the line in training?

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#20
post #7

Earlier quoted context omitted.

For textual purposes it seems fairly transformative. If you train a LLM on harry potter and ask it to generate a story that isn't harry potter then it's not a replacement. However, if you train a model on stock imagery and use it to generate stock imagery then I think you'll run into an issue from the Warhol case.

Wasn't that just over an arrangement of someone else's photographs?

https://en.wikipedia.org/wiki/Andy_Warhol_Foundation_for_the...

I wouldn't call it that. Goldsmith took a photograph of Prince which Warhol used as a reference to generate an illustration. Vanity Fair then chose to buy a license Warhol's print instead of Goldsmith's photograph.

So, despite the artwork being visual transformative (silkscreen vs photograph) the actual use was not transformed.

Post reply on HN