Live data from Hacker News

Thomson Reuters wins first major AI copyright case in the US

wired.com

141–150 of 188 posts

Re: Thomson Reuters wins first major AI copyright case in the US

#141
post #23

Here's the full decision, which (like most decisions!) is largely written to be legible to non-lawyers: https://storage.courtlistener.com/recap/gov.uscourts.ded.721... The core story seems to be: Westlaw writes and owns headnotes that help lawyers find legal cases about a particular topic. Ross paid people to translate those headnotes into new text, trained an AI on the translations, and used those to make a model th…

Westlaw's headnotes are primarily just snippets of the case with tags attached. They are really crappy. I hate them. Some lawyers love them.

Westlaw protects them because they are the "value add." Otherwise their business model is "take published decisions the court is legally bound to provide for free and sell it to you."

An LLM today could easily recreate the headnotes in a far superior manner from scratch with the right prompt. I don't even think hallucinations would factor in on such a small task that was well regulated, but you can always just asterisk the headnotes and put a disclaimer on them.

Re: Thomson Reuters wins first major AI copyright case in the US

#142
post #53

Interesting to note from this 2020 story (when ROSS shut down) that the company was founded in 2014 and went out of business in 2020: https://www.lawnext.com/2020/12/legal-research-company-ross-... The fact that it took until 2024 for the case to resolve shows how long the wheels of justice can take to turn!

Litigation takes forever. Especially when you factor COVID in. I'm still litigating cases from over a decade ago that are probably several years from resolution, just in the district court. Then you can spin through appeals courts for another five years. And that's civil.

Criminal, especially a death row case, can take 20+ years to exhaust every level of appellate review. In Illinois there are at least nine levels of review available to you without going through second rounds of review, state habeas, and collateral attacks like applications for clemency, pardons etc. If you're not paying for lawyers, expect each level to take around two years or more.

Re: Thomson Reuters wins first major AI copyright case in the US

#143
post #76

At the heart of this is a very greedy racket:- court reporters who 'own' the copyright to every word spoken by anyone in court that they transcribe to a transcript that they do not own the source to (judges/witnesses/lawyers/defendants in truth own it) They then milk huge fees for these transcripts and limit use/access/derivative works with huge fees. An AI verbatim transcriber would up end them, so that will be prev…

No, their work is valuable and they deserve to make money off of it.

The reason why it's valuable is it's transcribed live (usually with video) and is accurate and verifiable. Words and names are spelled correctly and speakers are correctly identified. Court reporters will stop speakers and ask for spelling or to repeat words.

AI transcriptions can't do that.

Re: Thomson Reuters wins first major AI copyright case in the US

#144

> Thomson Reuters prevailed on two of the four factors, but Bibas described the fourth as the most important, and ruled that Ross “meant to compete with Westlaw by developing a market substitute.” Yep. That's what people have been saying all along. If the intent is to substitute the original, then copying is not fair use. But the problem is that the current method for training requires this volume of data. So the mod…

>But the problem is that the current method for training requires this volume of data. So the models are legitimately not viable without massive copyright infringement.

Sure it is. It just requires what every other copyright'd work needs: permission and stipulations from the copyright holder. These aren't small time bloggers on the internet, these are large scale businesses.

>Though big-picture, it seems to me that the money-ed interests will ensure that even if the current legal landscape doesn't allow LLM's to exist, then they will lobby HARD until it is allowed.

The only solace I take is that these conglomerates are paying a lot to take down the rules they made 30 years ago when they weren't the ones profiting from stealing. But yes, I'm still frustrated by the hypocrisy.

Re: Thomson Reuters wins first major AI copyright case in the US

#145
post #16

Earlier quoted context omitted.

> the current method for training requires this volume of data This is one of those things that signal how dumb this technology still is - or maybe how smart humans are when compared to machines. A human brain doesn't need anywhere close to this volume of data, in order to be able to produce good output. I remember talking with friends 30 years ago about how it was inevitable that the brain would eventually be fully…

> A human brain doesn't need anywhere close to this volume of data, in order to be able to produce good output. Maybe not directly, but consider that our brains are the product of million of years of evolution and aren't a blank slate when we're born. Even though babies can't speak a language at birth, they already have all the neural connections in place in order to acquire and manipulate language, and require just…

sadly, those weights will not be inherited like they would to a baby. They'll be cooped up until the company dies, and that data probably dies with them. No wonder LLM has allegedly hit some stalls already.

Re: Thomson Reuters wins first major AI copyright case in the US

#146
post #136
post #14

Earlier quoted context omitted.

> So the models are legitimately not viable without massive copyright infringement. Copyright is not about acquisition, it is about publication and/or distribution. If I get a copy of Harry Potter from a dumpster, I can read it. If a company gets a copy of *all books from a torrent, they can use it to train their AI. The torrent providers may be in violation of copyright, and if the AI can be used to reproduce substa…

> Copyright is not about acquisition, it is about publication and/or distribution. It would be interesting to see how this holds up in court. "Your honor, I didn't watch the movie I downloaded, I only used it to train an AI." I highly suspect it would not matter.

well I think that will be the final judgement. We'll treat training data more as distribution than as consumption. Things always get more complicated when you put stuff up for sale. I also can't necessarily get away with Making "Garry Botter" who got accepted into an Enchanter school and goes on adventures with Jon and Germione. Unless it's parody, you can only cut so close before you're just infringinng anyway despite making it legally distinct.

Re: Thomson Reuters wins first major AI copyright case in the US

#147
I can't understand how some commenters frame such a result as not good. The big players will have no problem licensing large corpora to train their models, while my tiny site won't be vacuumed (legally at least) by scrapers if I won't agree.

My willingness to upload my projects anywhere is in the historical lows given the current state, honestly.

Re: Thomson Reuters wins first major AI copyright case in the US

#148

> Thomson Reuters prevailed on two of the four factors, but Bibas described the fourth as the most important, and ruled that Ross “meant to compete with Westlaw by developing a market substitute.” Yep. That's what people have been saying all along. If the intent is to substitute the original, then copying is not fair use. But the problem is that the current method for training requires this volume of data. So the mod…

>But the problem is that the current method for training requires this volume of data. So the models are legitimately not viable without massive copyright infringement. Sure it is. It just requires what every other copyright'd work needs: permission and stipulations from the copyright holder. These aren't small time bloggers on the internet, these are large scale businesses. >Though big-picture, it seems to me that t…

> Sure it is. It just requires what every other copyright'd work needs: permission and stipulations from the copyright holder.

Most other scenarios don't use millions/billions of works - that's the part which puts viability in question.

> these are large scale businesses.

I'd like training models to also remain accessible to open-source developers, academic researchers, and smaller businesses. Large-scale pretraining is common even for models that are not cutting-edge LLMs.

> The only solace I take is that these conglomerates are paying a lot to take down the rules they made 30 years ago when they weren't the ones profiting from stealing

As far as I'm aware, most of the lobbying in favor of stricter copyright has been done by Disney, Universal, Time Warner, RIAA, etc.

Not to say that tech companies have a consistent moral stance beyond whatever's currently in their financial self-interest, but I think that self-interest has put them in a position of supporting fair use and copyright safe harbors, opposing link tax, etc. more often than the the other way around - with cases like Authors Guild v. Google being a significant win for fair use.

Re: Thomson Reuters wins first major AI copyright case in the US

#149

Earlier quoted context omitted.

No that is not an extreme interpretation of the fair use factors. This is a routinely emphasized factor in fair use analyses for both copyright and trademark. School fair use is different because that defense is written into the statute directly in 17 U.S.C. § 107. Also, § 108 provides extensive protections for libraries and archives that go beyond fair use doctrines. The idea that the schools are encouraging the stu…

What a world we’re in where a school using text to teach children, who will remember it, talk about it with others, likely buy it for their own children… can be framed as a “massive indirect subsidy” rather than “free advertising”.

[deleted]

Re: Thomson Reuters wins first major AI copyright case in the US

#150
post #23

Here's the full decision, which (like most decisions!) is largely written to be legible to non-lawyers: https://storage.courtlistener.com/recap/gov.uscourts.ded.721... The core story seems to be: Westlaw writes and owns headnotes that help lawyers find legal cases about a particular topic. Ross paid people to translate those headnotes into new text, trained an AI on the translations, and used those to make a model th…

Westlaw's headnotes are primarily just snippets of the case with tags attached. They are really crappy. I hate them. Some lawyers love them. Westlaw protects them because they are the "value add." Otherwise their business model is "take published decisions the court is legally bound to provide for free and sell it to you." An LLM today could easily recreate the headnotes in a far superior manner from scratch with the…

Exactly. Why use the headnotes at all?

I always thought they were obviously were copyrightable. Plus they’re not close to perfect either.

Post reply on HN