Live data from Hacker News

Thomson Reuters wins first major AI copyright case in the US

wired.com

81–90 of 188 posts

Re: Thomson Reuters wins first major AI copyright case in the US

#81
post #16

> Thomson Reuters prevailed on two of the four factors, but Bibas described the fourth as the most important, and ruled that Ross “meant to compete with Westlaw by developing a market substitute.” Yep. That's what people have been saying all along. If the intent is to substitute the original, then copying is not fair use. But the problem is that the current method for training requires this volume of data. So the mod…

> the current method for training requires this volume of data This is one of those things that signal how dumb this technology still is - or maybe how smart humans are when compared to machines. A human brain doesn't need anywhere close to this volume of data, in order to be able to produce good output. I remember talking with friends 30 years ago about how it was inevitable that the brain would eventually be fully…

> I remember talking with friends 30 years ago about how it was inevitable that the brain would eventually be fully implemented as machine, once calculation power gets big enough; but it looks like we're still very far from that.

Why would it be?

"It's inevitable that the Burj Khalifa gets built, once steel production gets high enough."

"It's inevitable that Pegasuses will be bred from horses, as soon as somebody collects enough oats."

Reducing intelligence to the bulk aggregate of brute "calculation power" is... Ironically missing the point of intelligence.

Re: Thomson Reuters wins first major AI copyright case in the US

#82
post #14

Earlier quoted context omitted.

> So the models are legitimately not viable without massive copyright infringement. Copyright is not about acquisition, it is about publication and/or distribution. If I get a copy of Harry Potter from a dumpster, I can read it. If a company gets a copy of *all books from a torrent, they can use it to train their AI. The torrent providers may be in violation of copyright, and if the AI can be used to reproduce substa…

> simply training a model on illegally distributed text should not be copyright infringement You can train a model on copyrighted text, you just can't distribute the output in any way without violating copyright. (edit: depending on the other fair use factors). One of the big problems is that training is a mechanical process, so there is a direct line between the copyrighted works and the model's output, regardless o…

What’s a “mechanical process”? If I read The Lord of the Rings and it teaches me to write Star Wars, is that a mechanical process? My brain is governed by the laws of physics, right?

What if I’m a simulated brain running on a chip? What if I’m just a super-smart human and instead of reading and writing in the conventional way, I work out the LLM math in my head to generate the output?

Re: Thomson Reuters wins first major AI copyright case in the US

#85
post #23

Here's the full decision, which (like most decisions!) is largely written to be legible to non-lawyers: https://storage.courtlistener.com/recap/gov.uscourts.ded.721... The core story seems to be: Westlaw writes and owns headnotes that help lawyers find legal cases about a particular topic. Ross paid people to translate those headnotes into new text, trained an AI on the translations, and used those to make a model th…

This is an interesting opinion, but there are aspects of it that I doubt will stand the test of time. One aspect is the court’s ruling that West’s headnotes are copyrightable even when they merely quote a court opinion verbatim, because the editorial decision to quote the material itself shows a “creative spark”. It really isn’t workable — in law specifically - for copyright to attach to the mere selection of a quote…

The crux is Fair Use and until lobbyists change the four factor test, AI training has an uphill battle in court. It’s a very disliked observation in this forum, but I stand by my principles on this one because the courts see it my way. Derivative works, especially by artificial means, simply fail the test miserably and that’s the truth.

Re: Thomson Reuters wins first major AI copyright case in the US

#86
post #23

Here's the full decision, which (like most decisions!) is largely written to be legible to non-lawyers: https://storage.courtlistener.com/recap/gov.uscourts.ded.721... The core story seems to be: Westlaw writes and owns headnotes that help lawyers find legal cases about a particular topic. Ross paid people to translate those headnotes into new text, trained an AI on the translations, and used those to make a model th…

This is an interesting opinion, but there are aspects of it that I doubt will stand the test of time. One aspect is the court’s ruling that West’s headnotes are copyrightable even when they merely quote a court opinion verbatim, because the editorial decision to quote the material itself shows a “creative spark”. It really isn’t workable — in law specifically - for copyright to attach to the mere selection of a quote…

[deleted]

Re: Thomson Reuters wins first major AI copyright case in the US

#87
post #38

Thomson Reuters chose to sue Ross Intelligence, not a company like Google or even OpenAI. I wonder how deeper pockets would have affected the outcome. I wonder how the politics played out. The big AI companies could have funded Ross Intelligence, who could have threatened to sabotage their legal strategies by tanking and settling their own case in TR's favor.

I missed this line from the article:

Even before this ruling, Ross Intelligence had already felt the impact of the court battle: the startup shut down in 2021, citing the cost of litigation.

Re: Thomson Reuters wins first major AI copyright case in the US

#88
post #59

Earlier quoted context omitted.

How is the system inherently generative?

Generative is a technical term, meaning that a system models a full joint probability distribution. For example, a classifier is a generative model if it models p(example, label) -- which is sufficient to also calculate p(label | example) if you want -- rather than just modeling p(label | example) alone. Similar example in translation: a generative translation model would model p(french sentence, english sentence) --…

Your definition of "generative" as a statistical term is correct.

However, and annoyingly so, recently the general public and some experts have been speaking of "generative AI" (or GenAI for short) when they talk about large language models.

This creates the following contradiction:

- large language models are called "generative AI"

- large language models are based on transformers, which are neural networks

- neural networks are discriminative models (not generative ones like Hidden Markov Models)

- discriminative models are the oppositve of generative models, mathematically

So we may say "Generative AI is based on discriminative (not generative) classifiers and regressors". [as I am also a linguist, I regret this usage came into being, but in linguistics you describe how language is used, not how it should be used in a hypothetical world.]

References

- Gen AI (Wikipedia) https://en.wikipedia.org/wiki/Generative_artificial_intellig...

- Discriminative (Conditional) Model (Wikipedia) https://en.wikipedia.org/wiki/Discriminative_model

Re: Thomson Reuters wins first major AI copyright case in the US

#89

Earlier quoted context omitted.

License what? Every available copyrighted work? Even getting a tiny fraction is not practical. To the contrary, this just means companies can't make money from these models. Those using models for research and personal use wouldn't be infringing under the fair use tests.

> License what? Every available copyrighted work? Even getting a tiny fraction is not practical. They don't need every copyrighted work and getting a fraction is entirely practical. They would go to some large conglomerate like Getty Images or large publishers or social media whose terms give the site a license to what you post and then the middle men would get a vig and the original authors would get peanuts if anyt…

Applying copyright law more and more to things like software - and now to AI models - in other words, the status quo, makes little sense.

What is needed instead (I doubt politicians read HN, but someone go and tell them) is a new law that regulates training of these models if we want them to exist and be used in a legally safe way. This is needed for example because most jurisdictions have different copyright laws from one another, but software travels globally.

It would make sense to make all books available for non-commercial, perhaps even commercial R&D in AI, if society elected that to be beneficial in the same way that publishers must donate one copy of each new work to a copyright library (Library of Congress Library in the US, Oxford and Cambridge University libraries and British Library in the UK, Frankfurt and Leipzig Nationalbibliotheken for Germany etc.). Just add extra provisions that they need to send a plain text copy to the Linguistic Data Consortium (LDC), which manages datasets for NLP. Like for fair use, there can be provisions to make up for that use that happen automatically in the background (in some countries the price of photocopying machine includes a fee that gets passed on to copyright holders).

Otherwise you'll have one LLM being legal in one country but illegal in another because more than 15% from onw book were in the training data, and other messy situations.

Re: Thomson Reuters wins first major AI copyright case in the US

#90

How does this affect LLM systems that already have their corpus integrated?

The judge ruled this as a violation of copyright. Its the same as hosting any copyright material absent a valid license, criminal copyright piracy. They would need to figure out a way to prune the respective weights so that such material is not available, or risk legal fury.

> They would need to figure out a way to prune the respective weights so that such material is not available, or risk legal fury.

You want to reliably train it away from outputting the undesired outputs, not keeping it ignorant about them.

Post reply on HN