Live data from Hacker News

Thomson Reuters wins first major AI copyright case in the US

wired.com

131–140 of 188 posts

Re: Thomson Reuters wins first major AI copyright case in the US

#131
post #23

Here's the full decision, which (like most decisions!) is largely written to be legible to non-lawyers: https://storage.courtlistener.com/recap/gov.uscourts.ded.721... The core story seems to be: Westlaw writes and owns headnotes that help lawyers find legal cases about a particular topic. Ross paid people to translate those headnotes into new text, trained an AI on the translations, and used those to make a model th…

This is an interesting opinion, but there are aspects of it that I doubt will stand the test of time. One aspect is the court’s ruling that West’s headnotes are copyrightable even when they merely quote a court opinion verbatim, because the editorial decision to quote the material itself shows a “creative spark”. It really isn’t workable — in law specifically - for copyright to attach to the mere selection of a quote…

> It really isn’t workable — in law specifically - for copyright to attach to the mere selection of a quote from a case to represent that case’s holding on an issue. After all, we would expect many lawyers analyzing the case independently to converge on the same quotes!

I guess it depends on how long the source is, and how long the collection of quotes is, if we’d expect multiple lawyers to converge on the same solution. I don’t think it is totally obvious, though…

I’m also not sure if that’s a generally good test. It seems great for, like, painting. But I wouldn’t be surprised if we could come up with a photography scene where most professionals would converge on the same shot…

Re: Thomson Reuters wins first major AI copyright case in the US

#132
post #38

Thomson Reuters chose to sue Ross Intelligence, not a company like Google or even OpenAI. I wonder how deeper pockets would have affected the outcome. I wonder how the politics played out. The big AI companies could have funded Ross Intelligence, who could have threatened to sabotage their legal strategies by tanking and settling their own case in TR's favor.

It being in the legal realm probably had some impact. These tools can be seen as an attack on your profession, and I’m sure that affects the decision, whether conscious or not.

The judge is a user of Westlaw and not a shareholder, as is the judge's office. They would like Westlaw to be cheaper and easier to use. Westlaw takes the judge's work product and profits from it. Arguably, Westlaw should demand all judges recuse themselves!

Re: Thomson Reuters wins first major AI copyright case in the US

#133

Earlier quoted context omitted.

This is an interesting opinion, but there are aspects of it that I doubt will stand the test of time. One aspect is the court’s ruling that West’s headnotes are copyrightable even when they merely quote a court opinion verbatim, because the editorial decision to quote the material itself shows a “creative spark”. It really isn’t workable — in law specifically - for copyright to attach to the mere selection of a quote…

> That, plus the fact that Ross was a directly competing product, is what I see as really driving this decision. The "competing product" thing is probably the most extreme part of this opinion. The most important fair use factor is if the use competes with the original work, but this is generally implied to be directly competes, i.e. if you translate someone else's book from English to French and want to sell the tra…

I think that someone taking Biology 101 and ending up writing textbooks, as opposed to all the other people who just forgot what they learned once the elective was over or ended up working biologists with labs or teachers of biology and so forth, is quite different than someone saying hey I want to make a competing product to this successful company, let's take their content, re-write and use AI to make a competitor, and then actually going into direct competition with that company a couple years later

Re: Thomson Reuters wins first major AI copyright case in the US

#134
post #23

Here's the full decision, which (like most decisions!) is largely written to be legible to non-lawyers: https://storage.courtlistener.com/recap/gov.uscourts.ded.721... The core story seems to be: Westlaw writes and owns headnotes that help lawyers find legal cases about a particular topic. Ross paid people to translate those headnotes into new text, trained an AI on the translations, and used those to make a model th…

I'll quote a longer portion of the transcript about generative AI, because I think it makes the opposite of your point: Ross’s use is not transformative. Transformativeness is about the purpose of the use. “If an original work and a secondary use share the same or highly similar purposes, and the second use is of a commercial nature, the first factor is likely to weigh against fair use, absent some other justificatio…

I came here to point this out, and it's especially clear if you contextualize this with the original decision from September: https://www.ded.uscourts.gov/sites/ded/files/opinions/20-613...

They were doing semantic search using embeddings/rerankers.

The point that reading both decisions together compounds is that if they had trained a model on the Bulk Memos and generated novel text instead of doing direct searches, there likely would have been enough indirection introduced to prevent a summary judgement and this would have gone to a jury as the September decision states.

In other words, from their comment:

> But I'm not sure "generative" is that meaningful a distinction here.

The judge would not seem to agree at all.

Re: Thomson Reuters wins first major AI copyright case in the US

#135
post #29

This isn't really about "AI". It's about copying summaries. Google was fined for this in France for copying news headlines into their search results, and now has to pay royalties in the EU. Westlaw is a summarizing and indexing service for court case results. It's been publishing that info in book form since 1872. Ross was trying to compete with Westlaw, but used Westlaw as an input. West's "Key Numbers" are, after a…

The case involves headnotes, not just key numbers. Your links provide examples of such headnotes, which make it very clear that a lot of human creativity and judgment is involved in authoring them - they're not a matter of purely factual information, such as a phonebook. Thus, the headnotes are copywritten, and translating them to a different language doesn't negate that copyright. This looks like a slam dunk case, b…

> which make it very clear that a lot of human creativity and judgment is involved in authoring them

What's funny is that any SOTA LLM today could definitely author them, and even LexisNexis advertises the fact: https://www.lexisnexis.com/community/insights/legal/b/produc...

Re: Thomson Reuters wins first major AI copyright case in the US

#136
post #14

> Thomson Reuters prevailed on two of the four factors, but Bibas described the fourth as the most important, and ruled that Ross “meant to compete with Westlaw by developing a market substitute.” Yep. That's what people have been saying all along. If the intent is to substitute the original, then copying is not fair use. But the problem is that the current method for training requires this volume of data. So the mod…

> So the models are legitimately not viable without massive copyright infringement. Copyright is not about acquisition, it is about publication and/or distribution. If I get a copy of Harry Potter from a dumpster, I can read it. If a company gets a copy of *all books from a torrent, they can use it to train their AI. The torrent providers may be in violation of copyright, and if the AI can be used to reproduce substa…

> Copyright is not about acquisition, it is about publication and/or distribution.

It would be interesting to see how this holds up in court.

"Your honor, I didn't watch the movie I downloaded, I only used it to train an AI."

I highly suspect it would not matter.

Re: Thomson Reuters wins first major AI copyright case in the US

#137
post #41
post #29

This isn't really about "AI". It's about copying summaries. Google was fined for this in France for copying news headlines into their search results, and now has to pay royalties in the EU. Westlaw is a summarizing and indexing service for court case results. It's been publishing that info in book form since 1872. Ross was trying to compete with Westlaw, but used Westlaw as an input. West's "Key Numbers" are, after a…

TR may have intentionally chosen an easy battle to begin their legal war.

They began this case in 2020, before any of the most important models existed

Re: Thomson Reuters wins first major AI copyright case in the US

#138
post #63

Earlier quoted context omitted.

Open source model builders are no more entitled to rip off content owners than anyone else. I couldn't possibly care any less if this impacts "democratized access" to bullshit generators. At least if the big boys license the content then the rightful owners get paid (and have the option to opt out).

The copyright lobby has really done a number on public policy. Copyright was never meant to be perpetual. I’m good with your proposal if we also revert to the original 14 year + 14 year extension model. As it stands the 120 year copyright is so ridiculously tilted that we should not allow it to extend to veto power over technical advancements.

Legal arbitrage isn't a technical advancement. The technical advancement was all the stuff that goes into LLMs not the part where we feed ever more copyright into models for AICorp to make money.

Re: Thomson Reuters wins first major AI copyright case in the US

#139
post #23

Here's the full decision, which (like most decisions!) is largely written to be legible to non-lawyers: https://storage.courtlistener.com/recap/gov.uscourts.ded.721... The core story seems to be: Westlaw writes and owns headnotes that help lawyers find legal cases about a particular topic. Ross paid people to translate those headnotes into new text, trained an AI on the translations, and used those to make a model th…

> The court emphasizes "Because the AI landscape is changing rapidly, I note for readers that only non-generative AI is before me today." But I'm not sure "generative" is that meaningful a distinction here. Also the judge makes that statement, it looks like he misunderstands the nature of the AI system and the inherent generative elements it includes.

Does it really matter what the judge calls it when the ruling is about its end effects and outcomes?
Post reply on HN