Thomson Reuters wins first major AI copyright case in the US
21–30 of 188 posts
Re: Thomson Reuters wins first major AI copyright case in the US
#22Great decision for humans. Is this type of risk the reason why OpenAI masquerades as a non-profit?
Re: Thomson Reuters wins first major AI copyright case in the US
#23The core story seems to be: Westlaw writes and owns headnotes that help lawyers find legal cases about a particular topic. Ross paid people to translate those headnotes into new text, trained an AI on the translations, and used those to make a model that helps lawyers find legal cases about a particular topic. In that specific instance the court says this plan isn't fair use. If it was fair use, one could presumably just pay people to translate headnotes directly and make a Westlaw competitor, since translating headnotes is cheaper than writing new ones. And conversely if it isn't fair use where's the harm (the court notes no copyright violation was necessary for interoperability for example) -- one can still pay people to write fresh headnotes from caselaw and create the same training set.
The court emphasizes "Because the AI landscape is changing rapidly, I note for readers that only non-generative AI is before me today." But I'm not sure "generative" is that meaningful a distinction here.
You can definitely see how AI companies will be hustling to distinguish this from "we trained on copyrighted documents, and made a general purpose AI, and then people paid to use our AI to compete with the people who owned the documents." It's not quite the same, the connection is less direct, but it's not totally different.
Re: Thomson Reuters wins first major AI copyright case in the US
#24> Thomson Reuters prevailed on two of the four factors, but Bibas described the fourth as the most important, and ruled that Ross “meant to compete with Westlaw by developing a market substitute.” Yep. That's what people have been saying all along. If the intent is to substitute the original, then copying is not fair use. But the problem is that the current method for training requires this volume of data. So the mod…
> the current method for training requires this volume of data This is one of those things that signal how dumb this technology still is - or maybe how smart humans are when compared to machines. A human brain doesn't need anywhere close to this volume of data, in order to be able to produce good output. I remember talking with friends 30 years ago about how it was inevitable that the brain would eventually be fully…
> I remember talking with friends 30 years ago
I'd say you're pretty old. How many years of training did it take for you to start producing good output?
The leason here is we're kind of meta-trained: our minds are primed to pick up new things quickly by abstracting them and relating them to things we already know. We work in concepts and mental models rather than text. LLMs are incredibly weak by comparison. They only understand token sequences.
Re: Thomson Reuters wins first major AI copyright case in the US
#25> Thomson Reuters prevailed on two of the four factors, but Bibas described the fourth as the most important, and ruled that Ross “meant to compete with Westlaw by developing a market substitute.” Yep. That's what people have been saying all along. If the intent is to substitute the original, then copying is not fair use. But the problem is that the current method for training requires this volume of data. So the mod…
> So the models are legitimately not viable without massive copyright infringement. Copyright is not about acquisition, it is about publication and/or distribution. If I get a copy of Harry Potter from a dumpster, I can read it. If a company gets a copy of *all books from a torrent, they can use it to train their AI. The torrent providers may be in violation of copyright, and if the AI can be used to reproduce substa…
Re: Thomson Reuters wins first major AI copyright case in the US
#26> Thomson Reuters prevailed on two of the four factors, but Bibas described the fourth as the most important, and ruled that Ross “meant to compete with Westlaw by developing a market substitute.” Yep. That's what people have been saying all along. If the intent is to substitute the original, then copying is not fair use. But the problem is that the current method for training requires this volume of data. So the mod…
> So the models are legitimately not viable without massive copyright infringement. Copyright is not about acquisition, it is about publication and/or distribution. If I get a copy of Harry Potter from a dumpster, I can read it. If a company gets a copy of *all books from a torrent, they can use it to train their AI. The torrent providers may be in violation of copyright, and if the AI can be used to reproduce substa…
You can train a model on copyrighted text, you just can't distribute the output in any way without violating copyright. (edit: depending on the other fair use factors).
One of the big problems is that training is a mechanical process, so there is a direct line between the copyrighted works and the model's output, regardless of the form of the output. Just on those terms it is very likely to be a copyright violation. Even if they don't reproduce substantive portions, what they do reproduce is a derived work.
Re: Thomson Reuters wins first major AI copyright case in the US
#27> Thomson Reuters prevailed on two of the four factors, but Bibas described the fourth as the most important, and ruled that Ross “meant to compete with Westlaw by developing a market substitute.” Yep. That's what people have been saying all along. If the intent is to substitute the original, then copying is not fair use. But the problem is that the current method for training requires this volume of data. So the mod…
> the current method for training requires this volume of data This is one of those things that signal how dumb this technology still is - or maybe how smart humans are when compared to machines. A human brain doesn't need anywhere close to this volume of data, in order to be able to produce good output. I remember talking with friends 30 years ago about how it was inevitable that the brain would eventually be fully…
Maybe not directly, but consider that our brains are the product of million of years of evolution and aren't a blank slate when we're born. Even though babies can't speak a language at birth, they already have all the neural connections in place in order to acquire and manipulate language, and require just a few years of "supervised fine tuning" to learn the actual language.
LLMs, on the other hand, start with their weights at random values and need to catch up with those million years of evolution first.
Re: Thomson Reuters wins first major AI copyright case in the US
#28> Thomson Reuters prevailed on two of the four factors, but Bibas described the fourth as the most important, and ruled that Ross “meant to compete with Westlaw by developing a market substitute.” Yep. That's what people have been saying all along. If the intent is to substitute the original, then copying is not fair use. But the problem is that the current method for training requires this volume of data. So the mod…
> the current method for training requires this volume of data This is one of those things that signal how dumb this technology still is - or maybe how smart humans are when compared to machines. A human brain doesn't need anywhere close to this volume of data, in order to be able to produce good output. I remember talking with friends 30 years ago about how it was inevitable that the brain would eventually be fully…
Re: Thomson Reuters wins first major AI copyright case in the US
#29Ross was trying to compete with Westlaw, but used Westlaw as an input. West's "Key Numbers" are, after a century and a half, a de-facto standard.[2] So Ross had to match that proprietary indexing system to compete. Their output had to match Westlaw's rather closely. That's the underlying problem. The court ruled that the objective was to directly compete with Westlaw, and using Westlaw's output to do that was intentional copyright infringement.
This looks like a narrow holding, not one that generally covers feeding content into AI training systems.
[1] https://apnews.com/article/google-france-news-publishers-cop...
Re: Thomson Reuters wins first major AI copyright case in the US
#30> Thomson Reuters prevailed on two of the four factors, but Bibas described the fourth as the most important, and ruled that Ross “meant to compete with Westlaw by developing a market substitute.” Yep. That's what people have been saying all along. If the intent is to substitute the original, then copying is not fair use. But the problem is that the current method for training requires this volume of data. So the mod…
> So the models are legitimately not viable without massive copyright infringement. Copyright is not about acquisition, it is about publication and/or distribution. If I get a copy of Harry Potter from a dumpster, I can read it. If a company gets a copy of *all books from a torrent, they can use it to train their AI. The torrent providers may be in violation of copyright, and if the AI can be used to reproduce substa…
Simply running my business on illegally distributed copyrighted text/software/movie should not be copyright infringement.