Live data from Hacker News

Thomson Reuters wins first major AI copyright case in the US

wired.com

161–170 of 188 posts

Re: Thomson Reuters wins first major AI copyright case in the US

#161

Earlier quoted context omitted.

No that is not an extreme interpretation of the fair use factors. This is a routinely emphasized factor in fair use analyses for both copyright and trademark. School fair use is different because that defense is written into the statute directly in 17 U.S.C. § 107. Also, § 108 provides extensive protections for libraries and archives that go beyond fair use doctrines. The idea that the schools are encouraging the stu…

> School fair use is different because that defense is written into the statute directly It's written into the statute as an example of something that would be fair use. > The idea that the schools are encouraging the students to compete with the original authors of works taught in the classroom is fanciful by the meaning that courts usually apply to competition. People go to art school primarily because they want to…

>It's written into the statute as an example of something that would be fair use.

Statutory text controls what the courts can do, even and perhaps especially when it includes an example.

>People go to art school primarily because they want to create art. People study computer science primarily because they want to write code. It's their direct intention and purpose to compete with existing works.

Interesting perspective.

>So if you use Windows and then want to create Linux...

I don't understand your meaning.

>How is that logic any different than for AI training?

That is what Mark Lemley, law professor at Stanford, has argued in his many law review articles and amicus briefs: he believes that training is analogous to learning. The court here didn't agree with the Lemley view.

>It not only doesn't have any explicit requirement for a formal school (it just says "teaching"), it also isn't limited to teaching, teaching is just one of the things specified in the statute as being the kind of thing Congress intended fair use to include.

In practice courts tend to limit these exceptions to formal teaching arrangements.

Re: Thomson Reuters wins first major AI copyright case in the US

#162
Westlaw is to the legal profession what ResearchGate and others are to science research. They profit from information from the commons, and charge as much as the market will bear.

Only one of the many reasons the legal profession is so expensive.

Re: Thomson Reuters wins first major AI copyright case in the US

#163
post #23

Here's the full decision, which (like most decisions!) is largely written to be legible to non-lawyers: https://storage.courtlistener.com/recap/gov.uscourts.ded.721... The core story seems to be: Westlaw writes and owns headnotes that help lawyers find legal cases about a particular topic. Ross paid people to translate those headnotes into new text, trained an AI on the translations, and used those to make a model th…

AI has yet to demonstrate that it can do anything different from what a group of people could sit down and do. Sure, the AI may be able to do it faster, but there hasn't yet been anything demonstrated that exceeds what humans can do.

If it would be illegal for a group of people to do something, it is also going to be illegal for an AI do so.

Why is that so surprising?

Re: Thomson Reuters wins first major AI copyright case in the US

#164

Earlier quoted context omitted.

For that it would make more sense to run a routine which replaces letters with visually identical glyphs at different encoding points.

> For that it would make more sense to run a routine which replaces letters with visually identical glyphs at different encoding points. It seems like that would be pretty easily defeatable with the similar mapping to the one used to do the replacement.

Yes, but it's one more hurdle for them.

Re: Thomson Reuters wins first major AI copyright case in the US

#165

Earlier quoted context omitted.

If close paraphrase can be detected, this ought to be proof enough that some non-trivial element of creativity was involved in the original text. Because purely functional and necessary elements are not protected by copyright, even when they would otherwise be creative (this is technically known as the ' scenes à faire ' case) - and surely a "quote" which is unavoidable because it factually and unquestionably is the…

Isn't the argument that the act of selecting the right quote is the real work - and the work the copier avoided in the act of copying? You could argue that all the words are already in the dictionary - so none of them are new, you are just quoting from the dictionary in a particular order...... The reason you have people, rather than computers interpreting the law, is you can make judgements that make sense. Fundamen…

Copyright does not protect work ("sweat of the brow"), it only protects expression and creativity. Thus, whenever there is only one right expression or even a bare handful in any given context, copyright does not apply to that particular choice. By analogy, arranging words in some semi-arbitrary order can be an expressive choice, whereas using what's effectively a fixed phrase is not, even though the two might look similar and involve a comparable amount of "work".

Re: Thomson Reuters wins first major AI copyright case in the US

#166

Earlier quoted context omitted.

> That, plus the fact that Ross was a directly competing product, is what I see as really driving this decision. The "competing product" thing is probably the most extreme part of this opinion. The most important fair use factor is if the use competes with the original work, but this is generally implied to be directly competes, i.e. if you translate someone else's book from English to French and want to sell the tra…

The case looks pretty straightforward to me - they copied the notes ( human or machine doesn't really matter ) to directly compete with the author of the notes. If you wrote a program that automatically rephrased an original text - something like the Encyclopaedia Britannica - to preserve the meaning but not have identical phrasing - and then sold access to that information on in a way that undercut the original - th…

> automatically rephrased an original text - something like the Encyclopaedia Britannica - to preserve the meaning but not have identical phrasing

Note that it's very hard to do this starting from a single source, because in order to be safe from any copyright concern you'd have to only preserve the bare "idea" and everything else in your text must be independent. But LLM's seem to be able to get around this by looking at many sources that are all talking about the same facts and ideas in very different ways, and then successfully generalizing "out of sample" to a different expression of the same ideas.

Re: Thomson Reuters wins first major AI copyright case in the US

#167

Earlier quoted context omitted.

You're right as far as the MSJ is concerned, and I should've been more precise. I was focusing on the dictum in the preceding paragraph (because we're discussing the broader implications of the order rather than the nuts-and-bolts of the instant motion). In that paragraph, the judge wrote: > More than that, each headnote is an individual, copyrightable work. That became clear to me once I analogized the lawyer’s edit…

Yeah, I'm willing to bet that metaphor gets called out as ludicrous by a higher court, as it has broader implications across types of editorial expression that break down when examined. The marble from which a sculpture is carved is not itself a copyrighted work, and if we imagine it as having copyright protection, to the extent it's recognizable after editorial expression it'd have to qualify as fair use itself.

> Yeah, I'm willing to bet that metaphor gets called out as ludicrous by a higher court, as it has broader implications across types of editorial expression that break down when examined.

It's not ludicrous at all. Whether a work of "selection" from an existing source can be copyrightable in its own right would probably have to be judged on pretty much a case-by-case basis, but even in the context of "selecting" from a ruling there are almost certainly many cases where that work is creative and original enough that it can sensibly be protected by copyright.

Re: Thomson Reuters wins first major AI copyright case in the US

#168

Earlier quoted context omitted.

You're right as far as the MSJ is concerned, and I should've been more precise. I was focusing on the dictum in the preceding paragraph (because we're discussing the broader implications of the order rather than the nuts-and-bolts of the instant motion). In that paragraph, the judge wrote: > More than that, each headnote is an individual, copyrightable work. That became clear to me once I analogized the lawyer’s edit…

Yeah, I'm willing to bet that metaphor gets called out as ludicrous by a higher court, as it has broader implications across types of editorial expression that break down when examined. The marble from which a sculpture is carved is not itself a copyrighted work, and if we imagine it as having copyright protection, to the extent it's recognizable after editorial expression it'd have to qualify as fair use itself.

Both the more general premise (a work must not be an infringement of someone else’s work to be a work subject to copyright) and the more specific premise (court decisions are subject to copyright in the United States) in your argument for why verbatim selection from a court decision is not analogous, for copyright, to a sculptor carving from a block of material are wrong, though.

Re: Thomson Reuters wins first major AI copyright case in the US

#169
post #148

Earlier quoted context omitted.

>But the problem is that the current method for training requires this volume of data. So the models are legitimately not viable without massive copyright infringement. Sure it is. It just requires what every other copyright'd work needs: permission and stipulations from the copyright holder. These aren't small time bloggers on the internet, these are large scale businesses. >Though big-picture, it seems to me that t…

> Sure it is. It just requires what every other copyright'd work needs: permission and stipulations from the copyright holder. Most other scenarios don't use millions/billions of works - that's the part which puts viability in question. > these are large scale businesses. I'd like training models to also remain accessible to open-source developers, academic researchers, and smaller businesses. Large-scale pretraining…

>Most other scenarios don't use millions/billions of works - that's the part which puts viability in question.

Yes, they do. We have acquisitions in the billions these days and exclusivity deals in the hundreds of millions. Let's not pretend these companies can't do this through normal channels. They just wanna steal because they think they can get away from it.

>I'd like training models to also remain accessible to open-source developers, academic researchers, and smaller businesses.

Same. But such models still need to be ethically sourced. Maybe there's not enough royalty free content to compete with OpenAI, but it's pretty clear from Deepseek that you don't need 82 TB of data to be effective. If we need that much data, there are clearly optimizations to be made.

>I think that self-interest has put them in a position of supporting fair use and copyright safe harbors,

Yet they will sue anytime their data is scraped or otherwise not making the money. Maybe they didn't put trillions into lobbying like others, but they definitely have their fair share od using copyright. Microsoft won a lawsuit against web scraping via LinkedIn less than a year before OpenAI fell into legal troubles over scraping the entire internet.

Re: Thomson Reuters wins first major AI copyright case in the US

#170
post #53

Interesting to note from this 2020 story (when ROSS shut down) that the company was founded in 2014 and went out of business in 2020: https://www.lawnext.com/2020/12/legal-research-company-ross-... The fact that it took until 2024 for the case to resolve shows how long the wheels of justice can take to turn!

My father practiced corporate tax law and regularly had cases at trial that resolved issues from 20-30 years prior.

That's wild, I had no idea. I have trouble imagining a case where it's worth spending 30 years coming to a conclusion, but I guess that's one many reasons I'm not a corporate tax lawyer!
Post reply on HN