Earlier quoted context omitted.
For textual purposes it seems fairly transformative. If you train a LLM on harry potter and ask it to generate a story that isn't harry potter then it's not a replacement. However, if you train a model on stock imagery and use it to generate stock imagery then I think you'll run into an issue from the Warhol case.
The nature of how they store data makes it not okay in my books. You massage the data enough and you can generate something that seems infringement worthy.
A federal judge sides with Anthropic in lawsuit over training AI on books
61–70 of 222 posts
Re: A federal judge sides with Anthropic in lawsuit over training AI on books
#62Ok, so I can create a website, say, the-ai-pirate-bay.com, where I stream AI-reproduced movies. They are not verbatim, so I don't infringe any copyrights.
Re: A federal judge sides with Anthropic in lawsuit over training AI on books
#63Earlier quoted context omitted.
Indeed, I forsee a "training dataset consortium" arising out of this, whereby a bunch of companies team up to buy one copy of everything and then share it for training amongst themselves (ex. by reselling the entire library to each other for $1).
Like an Archive? Connected to the Internet?
Re: A federal judge sides with Anthropic in lawsuit over training AI on books
#64Earlier quoted context omitted.
But those training the LLMs are still using the works, and not just to discuss them, which I think is the point of fair use doctrine. I guess I fail to see how it's any different from me using it in some other way? If I wanted to write a play very loosely inspired by Blood Meridian, it might be transformative, but that doesn't justify me pirating the book. I tend to think copyright should be extremely limited compare…
> But those training the LLMs are still using the works, and not just to discuss them, which I think is the point of fair use doctrine. Worse, they’re using it for massive commercial gain, without paying a dime upstream to the supply chain that made it possible. If there is any purpose of copyright at all, it’s to prevent making money from someone’s else’s intellectual work. The entire thing is based on economic prag…
This makes no sense. If I buy and read a book on software engineering, and then use that knowledge to start a career, do I owe the author a percentage of my lifetime earnings?
Of course not. And yet I've made money with the help of someone else's intellectual work.
Copyright is actually pretty narrowly defined for _very good reason_.
Re: A federal judge sides with Anthropic in lawsuit over training AI on books
#65Earlier quoted context omitted.
The nature of how they store data makes it not okay in my books. You massage the data enough and you can generate something that seems infringement worthy.
For closed models the storage problem isn't really a problem, they can be judged by what they produce not how they store it as you don't have access to the actual data. That said, open weight LLMs are probably screwed, if enough of the work remains in the weights such that they can be extracted (even if it's without even talking to the LLM) then the weight file itself represents a copy of the work that's being distri…
Re: A federal judge sides with Anthropic in lawsuit over training AI on books
#66Earlier quoted context omitted.
For closed models the storage problem isn't really a problem, they can be judged by what they produce not how they store it as you don't have access to the actual data. That said, open weight LLMs are probably screwed, if enough of the work remains in the weights such that they can be extracted (even if it's without even talking to the LLM) then the weight file itself represents a copy of the work that's being distri…
Why doesn’t this apply to humans? If I memorize something such that it can be extracted did I violate the law? It’s only if I choose to allow such extraction to occur then I’m in violation of the law right? So if I or an LLM simply doesn’t allow said extraction to occur, memorization and copying is not against the law.
Re: A federal judge sides with Anthropic in lawsuit over training AI on books
#67Earlier quoted context omitted.
Why doesn’t this apply to humans? If I memorize something such that it can be extracted did I violate the law? It’s only if I choose to allow such extraction to occur then I’m in violation of the law right? So if I or an LLM simply doesn’t allow said extraction to occur, memorization and copying is not against the law.
I think an important distinction here is distribution... did you tell someone else what you memorized? Is downloading a model akin to distributing that same information?
Re: A federal judge sides with Anthropic in lawsuit over training AI on books
#68Earlier quoted context omitted.
Agreed. If I memorize a book and I am deployed into the world to talk about what I memorized that is not a violation of copyright. Which is reasonable logically because essentially this is what an LLM is doing.
It might be different if you are a commercial product which couldn’t have been created without incorporating the contents of all those books. Humans, animals, hardware and software are treated differently by law because they have different constraints and capabilities.
Let's be real, Humans have special treatment (more special than animals as we can eat and slaughter animals but not other humans) because WE created the law to serve humans.
So in terms of being fair across the board LLMs are no different. But there's no harm in giving ourselves special treatment.
Re: A federal judge sides with Anthropic in lawsuit over training AI on books
#69Earlier quoted context omitted.
I think an important distinction here is distribution... did you tell someone else what you memorized? Is downloading a model akin to distributing that same information?
What if I don't download the model and I just communicate with it. Sort of like chatting with another human. That's not a copyright issue right? I mean that's how most LLMs are deployed today.
Re: A federal judge sides with Anthropic in lawsuit over training AI on books
#70Earlier quoted context omitted.
> Definitely seems reasonable to say "you can train on this data but you have to have a legal copy" How many copies? They're not serving a single client. Libraries need to have multiple e-book licenses, after all.
In the human training case probably a Store DVD would still run afoul of that licensing issue. That's a broader topic of audience and I didn't want to muddy the analogy with that detail. It changes the definition of what a "legal copy" is but the general idea that the copy must be legal still stands.