Live data from Hacker News

Thomson Reuters wins first major AI copyright case in the US

wired.com

31–40 of 188 posts

Re: Thomson Reuters wins first major AI copyright case in the US

#31
post #8

Great. The stated goal of a lot of these companies seems to be “train the model on the output of humans, then hire us instead of the humans”. It’s been interesting that media where watermarking has been feasible (like photography) have seen creators get access to some compensation, while text based creators get nothing.

rotate similar [but different] fonts [or character pages] over each character. the sequence represents data thus watermark.

but the font changes won't be expressed in the (plain text) output of the LLM.

Re: Thomson Reuters wins first major AI copyright case in the US

#33
See. The fair-use excuses that the AI proponents here were trying to hang on to for dear life have fallen flat on this ruling.

This is going to be one of many cases in which there will be licensing deals being made out of this to stop AI grifters claiming 'fair use' to try to side-step copyright laws because they are using a gen AI system.

OpenAI ended up paying up for the data with Shutterstock and other news sources. This will be no different.

Re: Thomson Reuters wins first major AI copyright case in the US

#34
post #8

Earlier quoted context omitted.

rotate similar [but different] fonts [or character pages] over each character. the sequence represents data thus watermark.

but the font changes won't be expressed in the (plain text) output of the LLM.

Presumably the font will represent letters to look like a different letter, making it not useful to LLMs scraping the site but useful for visual readers.

This would have detrimental effects to people who use screen readers or have their own stylesheets of course.

Re: Thomson Reuters wins first major AI copyright case in the US

#35
post #16

Earlier quoted context omitted.

> the current method for training requires this volume of data This is one of those things that signal how dumb this technology still is - or maybe how smart humans are when compared to machines. A human brain doesn't need anywhere close to this volume of data, in order to be able to produce good output. I remember talking with friends 30 years ago about how it was inevitable that the brain would eventually be fully…

> A human brain doesn't need anywhere close to this volume of data, in order to be able to produce good output. Maybe not directly, but consider that our brains are the product of million of years of evolution and aren't a blank slate when we're born. Even though babies can't speak a language at birth, they already have all the neural connections in place in order to acquire and manipulate language, and require just…

Add to this, the brain is constantly processing raw sensory data from the moment it became viable, even when the body is "sleeping". It's using orders of magnitude more data than any model in existence every moment, but isn't generally deemed "intelligent" enough until it's around 18 years old.

Re: Thomson Reuters wins first major AI copyright case in the US

#36
post #14

> Thomson Reuters prevailed on two of the four factors, but Bibas described the fourth as the most important, and ruled that Ross “meant to compete with Westlaw by developing a market substitute.” Yep. That's what people have been saying all along. If the intent is to substitute the original, then copying is not fair use. But the problem is that the current method for training requires this volume of data. So the mod…

> So the models are legitimately not viable without massive copyright infringement. Copyright is not about acquisition, it is about publication and/or distribution. If I get a copy of Harry Potter from a dumpster, I can read it. If a company gets a copy of *all books from a torrent, they can use it to train their AI. The torrent providers may be in violation of copyright, and if the AI can be used to reproduce substa…

> Copyright is not about acquisition, it is about publication and/or distribution. If I get a copy of Harry Potter from a dumpster, I can read it. If a company gets a copy of *all books from a torrent, they can use it to train their AI.

"a person reading" and "computer processing of data" (training) are not the same thing

MDY Industries, LLC v. Blizzard Entertainment, Inc. rendered the verdict that loading unlicensed copyrighted material from disk was "copying", and hence copyright infringement

Re: Thomson Reuters wins first major AI copyright case in the US

#37

How does this affect LLM systems that already have their corpus integrated?

The judge ruled this as a violation of copyright. Its the same as hosting any copyright material absent a valid license, criminal copyright piracy. They would need to figure out a way to prune the respective weights so that such material is not available, or risk legal fury.

They would just censor output.

Youtube doesn't need to figure out how to stop copyright material from being uploaded, they need to stop it from being shared.

Re: Thomson Reuters wins first major AI copyright case in the US

#38
Thomson Reuters chose to sue Ross Intelligence, not a company like Google or even OpenAI. I wonder how deeper pockets would have affected the outcome.

I wonder how the politics played out. The big AI companies could have funded Ross Intelligence, who could have threatened to sabotage their legal strategies by tanking and settling their own case in TR's favor.

Re: Thomson Reuters wins first major AI copyright case in the US

#39
post #23

Here's the full decision, which (like most decisions!) is largely written to be legible to non-lawyers: https://storage.courtlistener.com/recap/gov.uscourts.ded.721... The core story seems to be: Westlaw writes and owns headnotes that help lawyers find legal cases about a particular topic. Ross paid people to translate those headnotes into new text, trained an AI on the translations, and used those to make a model th…

If the copyright holders win, the model giants will just license.

This effectively kills open source, which can't afford to license and won't be able to sublicense training data.

This is very bad for democratized access to and development of AI.

The giants will probably want this. The giants were already purchasing legacy media content enterprises (Amazon and MGM, etc.), so this will probably further consolidation and create extreme barriers to entry.

If I were OpenAI, I'd probably be very happy right now. If I were a recent batch YC AI company, I'd be mortified.

Re: Thomson Reuters wins first major AI copyright case in the US

#40
post #23

Here's the full decision, which (like most decisions!) is largely written to be legible to non-lawyers: https://storage.courtlistener.com/recap/gov.uscourts.ded.721... The core story seems to be: Westlaw writes and owns headnotes that help lawyers find legal cases about a particular topic. Ross paid people to translate those headnotes into new text, trained an AI on the translations, and used those to make a model th…

> The court emphasizes "Because the AI landscape is changing rapidly, I note for readers that only non-generative AI is before me today." But I'm not sure "generative" is that meaningful a distinction here.

Also the judge makes that statement, it looks like he misunderstands the nature of the AI system and the inherent generative elements it includes.

Post reply on HN