Live data from Hacker News

Thomson Reuters wins first major AI copyright case in the US

wired.com

171–180 of 188 posts

Re: Thomson Reuters wins first major AI copyright case in the US

#171
post #170

Earlier quoted context omitted.

My father practiced corporate tax law and regularly had cases at trial that resolved issues from 20-30 years prior.

That's wild, I had no idea. I have trouble imagining a case where it's worth spending 30 years coming to a conclusion, but I guess that's one many reasons I'm not a corporate tax lawyer!

One such repeating case ended up settling for over $10B, so it was definitely worth it!

To clarify, they spent decades litigating the same fundamental issue for each year’s tax filings, with each filing year taking multiple years to get to court. The plaintiffs won every single case until the government finally settled all the remaining tax years for that amount. Each year prior was worth hundreds of millions.

Re: Thomson Reuters wins first major AI copyright case in the US

#173
post #148

Earlier quoted context omitted.

> Sure it is. It just requires what every other copyright'd work needs: permission and stipulations from the copyright holder. Most other scenarios don't use millions/billions of works - that's the part which puts viability in question. > these are large scale businesses. I'd like training models to also remain accessible to open-source developers, academic researchers, and smaller businesses. Large-scale pretraining…

>Most other scenarios don't use millions/billions of works - that's the part which puts viability in question. Yes, they do. We have acquisitions in the billions these days and exclusivity deals in the hundreds of millions. Let's not pretend these companies can't do this through normal channels. They just wanna steal because they think they can get away from it. >I'd like training models to also remain accessible to…

> Yes, they do

To clarify: veggieroll said training models wouldn't be viable, you said it'd just require licensing like everyone else already manages, I said most other cases don't use millions/billions of works, you're saying that yes they do?

I feel like there must be a misunderstanding here, because that doesn't make much sense to me. Even for making a movie, which I think would be the most onerous of traditional cases, the number of works you'd license would likely be in the dozens (couple of pop songs, some stock images, etc.) - not billions.

> Let's not pretend these companies can't do this through normal channels

I'm not sure that there really has been a normal channel for licencing at the scale of "almost everything on the public Internet". A compulsory licensing scheme, like the US has for cover songs, could make it feasible to pay into a pot - but again I'd really hope for model training to remain accessible to smaller players opposed to just "meh, OpenAI has billions".

> but it's pretty clear from Deepseek that you don't need 82 TB of data to be effective.

As far as I'm aware, DeepSeek is not a low-data model. In fact, given China's more lax approach to copyright, I would not be surprised if the ability to freely pass around shadow libraries and large archives of torrented data without lawsuits was one of the factors contributing to their fast success relative to western counterparts.

> If we need that much data, there are clearly optimizations to be made.

I don't think this is necessarily a given - humans evolved on ~4 billion years worth of data, after all.

> Yet they will sue anytime their data is scraped or otherwise not making the money. Maybe they didn't put trillions into lobbying like others, but they definitely have their fair share od using copyright.

I believe lawsuits launched by or fuss kicked up by model developers will typically be on a contract basis (i.e "you agreed to our ToS then broke it") rather than a copyright basis. Again not to say these tech companies are acting in any way except their own self-interest, just that they've generally been more pro-fair-use than pro-strict-copyright on average to my knowledge.

Re: Thomson Reuters wins first major AI copyright case in the US

#174

Earlier quoted context omitted.

This is an interesting opinion, but there are aspects of it that I doubt will stand the test of time. One aspect is the court’s ruling that West’s headnotes are copyrightable even when they merely quote a court opinion verbatim, because the editorial decision to quote the material itself shows a “creative spark”. It really isn’t workable — in law specifically - for copyright to attach to the mere selection of a quote…

My experience using Westlaw Keycites at work is that they’re not primarily created by fishing a quote out of a holding, but instead by synthesizing a rule. If I want a summary, I read the Keycite; if I want a money quote, I root around in the case linked to the Keycite. Have you seen different? I’m curious what area of law you practice and in what state, for comparison’s sake.

Yeah, I'd agree that most are synthesized. But I do frequently see headnotes that are verbatim or nearly verbatim slices from the text. Just grabbing a case at random: Kearney v. Salomon Smith Barney, Inc., 39 Cal.4th 95 (2006). The 4th headnote reads:

> The federal system contemplates that individual states may adopt distinct policies to protect their own residents and generally may apply those policies to businesses that choose to conduct business within that state.

And the opinion reads:

> [T]he federal system contemplates that individual states may adopt distinct policies to protect their own residents and generally may apply those policies to businesses that choose to conduct business within that state.

Re: Thomson Reuters wins first major AI copyright case in the US

#175
post #23

Here's the full decision, which (like most decisions!) is largely written to be legible to non-lawyers: https://storage.courtlistener.com/recap/gov.uscourts.ded.721... The core story seems to be: Westlaw writes and owns headnotes that help lawyers find legal cases about a particular topic. Ross paid people to translate those headnotes into new text, trained an AI on the translations, and used those to make a model th…

> "we trained on copyrighted documents, and made a general purpose AI, and then people paid to use our AI to compete with the people who owned the documents"

This is a good distillation. A bit like "we trained our system on various works of art and music, and now it is being sold as a service that competes with the original artists and musicians."

Re: Thomson Reuters wins first major AI copyright case in the US

#176

Earlier quoted context omitted.

No that is not an extreme interpretation of the fair use factors. This is a routinely emphasized factor in fair use analyses for both copyright and trademark. School fair use is different because that defense is written into the statute directly in 17 U.S.C. § 107. Also, § 108 provides extensive protections for libraries and archives that go beyond fair use doctrines. The idea that the schools are encouraging the stu…

What a world we’re in where a school using text to teach children, who will remember it, talk about it with others, likely buy it for their own children… can be framed as a “massive indirect subsidy” rather than “free advertising”.

This reflects on the individuals choosing to create and proliferate such misleading or hyperbolic framing more than it does on the world that we all live in. In meatspace we usually reject these ideas and ignore the people pushing them.

Re: Thomson Reuters wins first major AI copyright case in the US

#177
post #173

Earlier quoted context omitted.

>Most other scenarios don't use millions/billions of works - that's the part which puts viability in question. Yes, they do. We have acquisitions in the billions these days and exclusivity deals in the hundreds of millions. Let's not pretend these companies can't do this through normal channels. They just wanna steal because they think they can get away from it. >I'd like training models to also remain accessible to…

> Yes, they do To clarify: veggieroll said training models wouldn't be viable, you said it'd just require licensing like everyone else already manages, I said most other cases don't use millions/billions of works, you're saying that yes they do? I feel like there must be a misunderstanding here, because that doesn't make much sense to me. Even for making a movie, which I think would be the most onerous of traditional…

> I said most other cases don't use millions/billions of works, you're saying that yes they do?

I assumed we were talking about logistics, not tech. I'm sure it will be technically possibly to use less training data overtime (Deepseek is more or less demonstrating that in real time. Maybe there's copyright data but I'd be surprised if it used anything close to 80 TB like competittorz).

I know hindsight is 20/20, but I always felt the earlier approaches were absurdly brute forced.

>I'm not sure that there really has been a normal channel for licencing at the scale of "almost everything on the public Internet"

There isn't. So they'd need to do it the old fashioned way with agreements . Or make some incentive model that has media submit their works with that understanding of training. Or any number of marketing ideas.

I don't exactly pity their herculean effort. Those same companies spend decades suing individuals for much pettier uses and building those precedent up (some covered under free use).

>and large archives of torrented data without lawsuits was one of the factors contributing to their fast success relative to western counterparts.

And now they're being slowed down. If not litigsted out of the market. Public trust in AI is falling. The lack of oversight into hallucinations may have even cost a few lives. Content creators now need to take extra precautions so they aren't stolen from because they don't even bother trying to respect robots.txt. Even a few posts here on HN note how the scraping is so rampant that it can spike their hosting costs on websites (so now we need more capthas. And I hate myself for uttering such a sentence).

Was all that velocity worth it? Who benefitted from this outside of a few billipnaires? We can't even say we beat China on this.

>I don't think this is necessarily a given - humans evolved on ~4 billion years worth of data, after all

Humans inherit their data and slowly structure around that. Maybe if AI models collaborated together as humanity did, I would sympathize more with this argument.

We both know it's instead a rat race and the goal isn't survival and passing on knowledge (and genes) to the next generation. AI can evolve organically but it instead devolved into a thieve's den.

I take the approach more like Bell's Spacecraft paradox. If they started gaining data ethically, by the time they gather a decent chunk they probably would have already optimized a model that needs less data. It'd be slower but not actually much slower I'm the long run. But they aren't exactly trying to go for quality here.

>I believe lawsuits launched by or fuss kicked up by model developers will typically be on a contract basis (i.e "you agreed to our ToS then broke it") rather than a copyright basis.

I suppose we'll see. Too early to tell. This lawsuit will definitely be precedent in other ongoing cases, but others may shift to a copyright infringement case anyway. Unlike other llms there was some human tailoring going on here, so it's not fully comparable to something like the NYT case.

Re: Thomson Reuters wins first major AI copyright case in the US

#178

Earlier quoted context omitted.

Isn't the argument that the act of selecting the right quote is the real work - and the work the copier avoided in the act of copying? You could argue that all the words are already in the dictionary - so none of them are new, you are just quoting from the dictionary in a particular order...... The reason you have people, rather than computers interpreting the law, is you can make judgements that make sense. Fundamen…

Copyright does not protect work ("sweat of the brow"), it only protects expression and creativity. Thus, whenever there is only one right expression or even a bare handful in any given context, copyright does not apply to that particular choice. By analogy, arranging words in some semi-arbitrary order can be an expressive choice, whereas using what's effectively a fixed phrase is not, even though the two might look s…

The intention of copyright is to protect useful work.

The detail of how to do that in fair way that doesn't block other people is complex[1] - you can never cover all possibilities in a written law - that's why you have people interpreting them and making judgements. All I'm saying is the guiding light in that interpretation is copyright is there to protect the justifiable work of people in a fair way.

Somebody taking those law notes and trivially copying them to directly compete is clearly not 'fair use'.

If those notes could have been created mechanically directly from the original source - why didn't the copier do that - rather than use the competitors work?

[1] given the endless creativity of humans to game systems.

Re: Thomson Reuters wins first major AI copyright case in the US

#179

Earlier quoted context omitted.

The case looks pretty straightforward to me - they copied the notes ( human or machine doesn't really matter ) to directly compete with the author of the notes. If you wrote a program that automatically rephrased an original text - something like the Encyclopaedia Britannica - to preserve the meaning but not have identical phrasing - and then sold access to that information on in a way that undercut the original - th…

> automatically rephrased an original text - something like the Encyclopaedia Britannica - to preserve the meaning but not have identical phrasing Note that it's very hard to do this starting from a single source, because in order to be safe from any copyright concern you'd have to only preserve the bare "idea" and everything else in your text must be independent. But LLM's seem to be able to get around this by looki…

The concept clustering across multiple sources allows you to rephrase more accurately while retaining meaning - however the point I'm making is if you then point that program at Encyclopaedia Britannica and simply rephrase it then charge for access to the rephrased version - should you be allowed to do that?

Re: Thomson Reuters wins first major AI copyright case in the US

#180

Earlier quoted context omitted.

Copyright does not protect work ("sweat of the brow"), it only protects expression and creativity. Thus, whenever there is only one right expression or even a bare handful in any given context, copyright does not apply to that particular choice. By analogy, arranging words in some semi-arbitrary order can be an expressive choice, whereas using what's effectively a fixed phrase is not, even though the two might look s…

The intention of copyright is to protect useful work. The detail of how to do that in fair way that doesn't block other people is complex[1] - you can never cover all possibilities in a written law - that's why you have people interpreting them and making judgements . All I'm saying is the guiding light in that interpretation is copyright is there to protect the justifiable work of people in a fair way. Somebody taki…

> The intention of copyright is

..."to promote the progress of science and useful arts". I don't see anything in there about rewarding 'work' irrespective of whether that work involves any kind of creativity.

> If those notes could have been created mechanically directly from the original source - why didn't the copier do that

That's actually a very good question. In practice, I do absolutely agree that the notes involve plenty of originality and creativity.

Post reply on HN