Live data from Hacker News

Thomson Reuters wins first major AI copyright case in the US

wired.com

111–120 of 188 posts

Re: Thomson Reuters wins first major AI copyright case in the US

#111
post #60
post #52

Earlier quoted context omitted.

"The biggest models want to train on literally every piece of human-written text ever written" They genuinely don't. There is a LOT of garbage text out there that they don't want. They want to train on every high quality piece of human-written text they can get their hands on (where the definition of "high quality" is a major piece of the secret sauce that makes some LLMs better than others), but that doesn't mean ev…

Even restricted to that narrower definition, the major commercial model companies wouldn't be able to afford to license all their high-quality human text. OpenAI is Uber with a slightly less ethically despicable CEO. It knows it's flaunting the spirit of copyright law -- it's just hoping it could bootstrap quickly enough to make the question irrelevant. If every commercial AI company that couldn't prove training data…

Bold idea, requiring startups to proactively prove they have not broken the law. Should we apply it to all tech startups? Let’s see silicon startups prove they have not stolen trade secrets!

Re: Thomson Reuters wins first major AI copyright case in the US

#112
post #23

Here's the full decision, which (like most decisions!) is largely written to be legible to non-lawyers: https://storage.courtlistener.com/recap/gov.uscourts.ded.721... The core story seems to be: Westlaw writes and owns headnotes that help lawyers find legal cases about a particular topic. Ross paid people to translate those headnotes into new text, trained an AI on the translations, and used those to make a model th…

> You can definitely see how AI companies will be hustling to distinguish this from "we trained on copyrighted documents, and made a general purpose AI, and then people paid to use our AI to compete with the people who owned the documents." It's not quite the same, the connection is less direct, but it's not totally different. Surely creating a general-purpose AI is transformative, though? Are you anticipating that A…

IMO yes. The entire purpose of copyright law is to protect the incentive to create new material. A huge portion of the value prop of AI is that it captures the incentive normally bound for the creators of the training material (i.e. the whole point is you can ask the AI and not even see, never mind pay, the originator).

Re: Thomson Reuters wins first major AI copyright case in the US

#113

Earlier quoted context omitted.

No that is not an extreme interpretation of the fair use factors. This is a routinely emphasized factor in fair use analyses for both copyright and trademark. School fair use is different because that defense is written into the statute directly in 17 U.S.C. § 107. Also, § 108 provides extensive protections for libraries and archives that go beyond fair use doctrines. The idea that the schools are encouraging the stu…

> School fair use is different because that defense is written into the statute directly It's written into the statute as an example of something that would be fair use. > The idea that the schools are encouraging the students to compete with the original authors of works taught in the classroom is fanciful by the meaning that courts usually apply to competition. People go to art school primarily because they want to…

Copyright covers expression, not ideas. The underlying problem here is that Ross Intelligence never went to the trouble of distilling the purely idea-based and factual element from their original sources; even their finalized search system still had a pervasive reliance on Westlaw's original and creative expression as embedded in their headnotes. Using Windows and then creating Linux is something entirely different because Linux goes to great effort in order not to use anything that's specific to Windows. Large-scale language models are probably somewhere in the middle, because their unique reliance on an incredibly wide variety of published texts makes it very unlikely that they'll ever preserve anything of substance about the expression in any single text.

Re: Thomson Reuters wins first major AI copyright case in the US

#114
post #23

Here's the full decision, which (like most decisions!) is largely written to be legible to non-lawyers: https://storage.courtlistener.com/recap/gov.uscourts.ded.721... The core story seems to be: Westlaw writes and owns headnotes that help lawyers find legal cases about a particular topic. Ross paid people to translate those headnotes into new text, trained an AI on the translations, and used those to make a model th…

This is an interesting opinion, but there are aspects of it that I doubt will stand the test of time. One aspect is the court’s ruling that West’s headnotes are copyrightable even when they merely quote a court opinion verbatim, because the editorial decision to quote the material itself shows a “creative spark”. It really isn’t workable — in law specifically - for copyright to attach to the mere selection of a quote…

>the court’s ruling that West’s headnotes are copyrightable even when they merely quote a court opinion verbatim, because the editorial decision to quote the material itself shows a “creative spark” ... when Ross paid human annotators to write their own versions of the headnotes, they really did crib from West’s wholesale rather than doing their own independent analysis

... so it follows that it was then Ross's annotators showing the creative spark

Re: Thomson Reuters wins first major AI copyright case in the US

#115
post #63
post #39

Earlier quoted context omitted.

If the copyright holders win, the model giants will just license. This effectively kills open source, which can't afford to license and won't be able to sublicense training data. This is very bad for democratized access to and development of AI. The giants will probably want this. The giants were already purchasing legacy media content enterprises (Amazon and MGM, etc.), so this will probably further consolidation an…

Open source model builders are no more entitled to rip off content owners than anyone else. I couldn't possibly care any less if this impacts "democratized access" to bullshit generators. At least if the big boys license the content then the rightful owners get paid (and have the option to opt out).

The copyright lobby has really done a number on public policy. Copyright was never meant to be perpetual.

I’m good with your proposal if we also revert to the original 14 year + 14 year extension model. As it stands the 120 year copyright is so ridiculously tilted that we should not allow it to extend to veto power over technical advancements.

Re: Thomson Reuters wins first major AI copyright case in the US

#116

Earlier quoted context omitted.

> You can definitely see how AI companies will be hustling to distinguish this from "we trained on copyrighted documents, and made a general purpose AI, and then people paid to use our AI to compete with the people who owned the documents." It's not quite the same, the connection is less direct, but it's not totally different. Surely creating a general-purpose AI is transformative, though? Are you anticipating that A…

IMO yes. The entire purpose of copyright law is to protect the incentive to create new material. A huge portion of the value prop of AI is that it captures the incentive normally bound for the creators of the training material (i.e. the whole point is you can ask the AI and not even see, never mind pay, the originator).

I'm not a lawyer, but I think the bar for contributory infringement is much higher than that. I think you'd have to find representatives of the defendants actually indicating somehow that people should use it that way. It seems to me that Grokster, etc.'s encouragement of their users to infringe copyright was an important factor in them losing this case, for instance.

https://supreme.justia.com/cases/federal/us/545/913/

Re: Thomson Reuters wins first major AI copyright case in the US

#117
Almost every article I read on fair use talked like I could only use small amounts while not competing with them. AI people focus on a tiny number of precedents that they stretch very far. A reasonable person wouldn’t come up with their interpretation of fair use after looking at how most examples play out in court.

It shouldn’t surprise the writer that the AI companies’ versions of fair use didn’t hold much weight. They should assume that would be true. Then, be surprised any time a pro-AI ruling goes against common examples in case law. The AI companies are hoping to achieve that by throwing enough money at the legal system.

Re: Thomson Reuters wins first major AI copyright case in the US

#118

Earlier quoted context omitted.

> You can definitely see how AI companies will be hustling to distinguish this from "we trained on copyrighted documents, and made a general purpose AI, and then people paid to use our AI to compete with the people who owned the documents." It's not quite the same, the connection is less direct, but it's not totally different. Surely creating a general-purpose AI is transformative, though? Are you anticipating that A…

IMO yes. The entire purpose of copyright law is to protect the incentive to create new material. A huge portion of the value prop of AI is that it captures the incentive normally bound for the creators of the training material (i.e. the whole point is you can ask the AI and not even see, never mind pay, the originator).

Ask the AI for what exactly? Factual information? That gets very low protection from a copyright point of view, especially when separate random answers by the AI will routinely show completely different rephrasings of the AI's response - implying that it can generalize well beyond the "expression" contained in any single answer, and effectively reference the underlying facts.

Re: Thomson Reuters wins first major AI copyright case in the US

#119

Earlier quoted context omitted.

IMO yes. The entire purpose of copyright law is to protect the incentive to create new material. A huge portion of the value prop of AI is that it captures the incentive normally bound for the creators of the training material (i.e. the whole point is you can ask the AI and not even see, never mind pay, the originator).

I'm not a lawyer, but I think the bar for contributory infringement is much higher than that. I think you'd have to find representatives of the defendants actually indicating somehow that people should use it that way. It seems to me that Grokster, etc.'s encouragement of their users to infringe copyright was an important factor in them losing this case, for instance. https://supreme.justia.com/cases/federal/us/545/9…

Encouragement is definitely not a required element. https://www.cantorcolburn.com/news-newsletters-387.html

Re: Thomson Reuters wins first major AI copyright case in the US

#120
post #29

This isn't really about "AI". It's about copying summaries. Google was fined for this in France for copying news headlines into their search results, and now has to pay royalties in the EU. Westlaw is a summarizing and indexing service for court case results. It's been publishing that info in book form since 1872. Ross was trying to compete with Westlaw, but used Westlaw as an input. West's "Key Numbers" are, after a…

The case involves headnotes, not just key numbers. Your links provide examples of such headnotes, which make it very clear that a lot of human creativity and judgment is involved in authoring them - they're not a matter of purely factual information, such as a phonebook. Thus, the headnotes are copywritten, and translating them to a different language doesn't negate that copyright. This looks like a slam dunk case, b…

Case law is public domain. You can publish digitized copies of Westlaw books with the headnotes, keys, and a couple of other property bits redacted. Any of their proprietary elements though, definitely including the key cites, are clearly a no-go. The headnotes not only require creativity and expertise to make, many lawyers consider them indispensable (though many other lawyers apparently throw shade at lawyers that rely on them.) And since most of the rest of the book is public domain, it’s one of the biggest, if not the biggest selling point for their texts. They famously vigorously defend their copyrights— the defendant surely knew what they were signing up for when they started doing this.
Post reply on HN