Live data from Hacker News

Thomson Reuters wins first major AI copyright case in the US

wired.com

11–20 of 188 posts

Re: Thomson Reuters wins first major AI copyright case in the US

#13

How does this affect LLM systems that already have their corpus integrated?

The judge ruled this as a violation of copyright. Its the same as hosting any copyright material absent a valid license, criminal copyright piracy.

They would need to figure out a way to prune the respective weights so that such material is not available, or risk legal fury.

Re: Thomson Reuters wins first major AI copyright case in the US

#14

> Thomson Reuters prevailed on two of the four factors, but Bibas described the fourth as the most important, and ruled that Ross “meant to compete with Westlaw by developing a market substitute.” Yep. That's what people have been saying all along. If the intent is to substitute the original, then copying is not fair use. But the problem is that the current method for training requires this volume of data. So the mod…

> So the models are legitimately not viable without massive copyright infringement.

Copyright is not about acquisition, it is about publication and/or distribution. If I get a copy of Harry Potter from a dumpster, I can read it. If a company gets a copy of *all books from a torrent, they can use it to train their AI. The torrent providers may be in violation of copyright, and if the AI can be used to reproduce substantive portions of the original text, the AI companies then may be in violation of copyright, but simply training a model on illegally distributed text should not be copyright infringement.

Re: Thomson Reuters wins first major AI copyright case in the US

#15
post #14

> Thomson Reuters prevailed on two of the four factors, but Bibas described the fourth as the most important, and ruled that Ross “meant to compete with Westlaw by developing a market substitute.” Yep. That's what people have been saying all along. If the intent is to substitute the original, then copying is not fair use. But the problem is that the current method for training requires this volume of data. So the mod…

> So the models are legitimately not viable without massive copyright infringement. Copyright is not about acquisition, it is about publication and/or distribution. If I get a copy of Harry Potter from a dumpster, I can read it. If a company gets a copy of *all books from a torrent, they can use it to train their AI. The torrent providers may be in violation of copyright, and if the AI can be used to reproduce substa…

I mean, you're right in the abstract. If you train an LLM in a void and never do anything with the model, sure.

But that's not what anyone is doing. People train models so that someone can actually use them. So I'm not sure how your comment is helpful other than to point out that distinction (which doesn't make much difference in this case specifically or how copyright applies for LLM's in general)

Re: Thomson Reuters wins first major AI copyright case in the US

#16

> Thomson Reuters prevailed on two of the four factors, but Bibas described the fourth as the most important, and ruled that Ross “meant to compete with Westlaw by developing a market substitute.” Yep. That's what people have been saying all along. If the intent is to substitute the original, then copying is not fair use. But the problem is that the current method for training requires this volume of data. So the mod…

> the current method for training requires this volume of data

This is one of those things that signal how dumb this technology still is - or maybe how smart humans are when compared to machines. A human brain doesn't need anywhere close to this volume of data, in order to be able to produce good output.

I remember talking with friends 30 years ago about how it was inevitable that the brain would eventually be fully implemented as machine, once calculation power gets big enough; but it looks like we're still very far from that.

Re: Thomson Reuters wins first major AI copyright case in the US

#17
post #16

> Thomson Reuters prevailed on two of the four factors, but Bibas described the fourth as the most important, and ruled that Ross “meant to compete with Westlaw by developing a market substitute.” Yep. That's what people have been saying all along. If the intent is to substitute the original, then copying is not fair use. But the problem is that the current method for training requires this volume of data. So the mod…

> the current method for training requires this volume of data This is one of those things that signal how dumb this technology still is - or maybe how smart humans are when compared to machines. A human brain doesn't need anywhere close to this volume of data, in order to be able to produce good output. I remember talking with friends 30 years ago about how it was inevitable that the brain would eventually be fully…

[deleted]

Re: Thomson Reuters wins first major AI copyright case in the US

#19
post #14

> Thomson Reuters prevailed on two of the four factors, but Bibas described the fourth as the most important, and ruled that Ross “meant to compete with Westlaw by developing a market substitute.” Yep. That's what people have been saying all along. If the intent is to substitute the original, then copying is not fair use. But the problem is that the current method for training requires this volume of data. So the mod…

> So the models are legitimately not viable without massive copyright infringement. Copyright is not about acquisition, it is about publication and/or distribution. If I get a copy of Harry Potter from a dumpster, I can read it. If a company gets a copy of *all books from a torrent, they can use it to train their AI. The torrent providers may be in violation of copyright, and if the AI can be used to reproduce substa…

[deleted]

Re: Thomson Reuters wins first major AI copyright case in the US

#20
post #16

> Thomson Reuters prevailed on two of the four factors, but Bibas described the fourth as the most important, and ruled that Ross “meant to compete with Westlaw by developing a market substitute.” Yep. That's what people have been saying all along. If the intent is to substitute the original, then copying is not fair use. But the problem is that the current method for training requires this volume of data. So the mod…

> the current method for training requires this volume of data This is one of those things that signal how dumb this technology still is - or maybe how smart humans are when compared to machines. A human brain doesn't need anywhere close to this volume of data, in order to be able to produce good output. I remember talking with friends 30 years ago about how it was inevitable that the brain would eventually be fully…

Absolutely! But the question is whether the next step-change in intelligence is just around the corner (in which case, this legal speedbump might spur innovation). Or, will the next revolution take a while.

There's enough money in the market to fund a lot of research into totally novel underlying methods. But if it takes too long, investors and lawmakers will just move to make what already works legal, because it is useful.

Post reply on HN