Live data from Hacker News

Thomson Reuters wins first major AI copyright case in the US

wired.com

121–130 of 188 posts

Re: Thomson Reuters wins first major AI copyright case in the US

#121
post #34

Earlier quoted context omitted.

but the font changes won't be expressed in the (plain text) output of the LLM.

Presumably the font will represent letters to look like a different letter, making it not useful to LLMs scraping the site but useful for visual readers. This would have detrimental effects to people who use screen readers or have their own stylesheets of course.

That would only stymie the smallest-time players. Things like sideways text in margins or rotated table column headers are common enough that these have been solved problems for decades. Breaking the text down into specific elements and handling it differently or ignoring it altogether based on context and content is trivial.

Re: Thomson Reuters wins first major AI copyright case in the US

#122
post #47
post #33

See. The fair-use excuses that the AI proponents here were trying to hang on to for dear life have fallen flat on this ruling. This is going to be one of many cases in which there will be licensing deals being made out of this to stop AI grifters claiming 'fair use' to try to side-step copyright laws because they are using a gen AI system. OpenAI ended up paying up for the data with Shutterstock and other news source…

My biggest concern is, what happens when countries like China, who aren't restricted by this, far outpace western countries in this technology? Do we just shrug and accept our far inferior models? LLMs are a productivity multiplier (similar to a search engine), so it'll have a large impact on the economy if licensing costs prohibit large scale training.

"Our" models are not inferior. There is plenty of data, and the next frontier is prediction-time compute and data synthesis. Shouldn't the Chinese worry that they are depressing the publication of commercial IP?

Re: Thomson Reuters wins first major AI copyright case in the US

#123
post #38

Thomson Reuters chose to sue Ross Intelligence, not a company like Google or even OpenAI. I wonder how deeper pockets would have affected the outcome. I wonder how the politics played out. The big AI companies could have funded Ross Intelligence, who could have threatened to sabotage their legal strategies by tanking and settling their own case in TR's favor.

It being in the legal realm probably had some impact. These tools can be seen as an attack on your profession, and I’m sure that affects the decision, whether conscious or not.

Re: Thomson Reuters wins first major AI copyright case in the US

#124
post #14

Earlier quoted context omitted.

> So the models are legitimately not viable without massive copyright infringement. Copyright is not about acquisition, it is about publication and/or distribution. If I get a copy of Harry Potter from a dumpster, I can read it. If a company gets a copy of *all books from a torrent, they can use it to train their AI. The torrent providers may be in violation of copyright, and if the AI can be used to reproduce substa…

That's an interesting take, but false in a lot of juristictions. Even if we ignore question of if the model can distribute work, in many places even downloading content is illegal. Otherwise the person torrenting a movie would be totally in the clear, or thing about what MS would say if a company "just" downloads copies of Windows to use on their computers without ever distributing them.

>Otherwise the person torrenting a movie would be totally in the clear

Any examples of people being sued for merely downloading? "Torrenting" basically always involves uploading, even if you stop immediately after completion. A better test would be if someone was sued for using an illegal streaming site, which to my knowledge has never happened.

Re: Thomson Reuters wins first major AI copyright case in the US

#125
post #14

Earlier quoted context omitted.

> So the models are legitimately not viable without massive copyright infringement. Copyright is not about acquisition, it is about publication and/or distribution. If I get a copy of Harry Potter from a dumpster, I can read it. If a company gets a copy of *all books from a torrent, they can use it to train their AI. The torrent providers may be in violation of copyright, and if the AI can be used to reproduce substa…

> simply training a model on illegally distributed text should not be copyright infringement You can train a model on copyrighted text, you just can't distribute the output in any way without violating copyright. (edit: depending on the other fair use factors). One of the big problems is that training is a mechanical process, so there is a direct line between the copyrighted works and the model's output, regardless o…

>One of the big problems is that training is a mechanical process, so there is a direct line between the copyrighted works and the model's output, regardless of the form of the output. Just on those terms it is very likely to be a copyright violation. Even if they don't reproduce substantive portions, what they do reproduce is a derived work.

Google making thumbnails or scanning books are both arguably "mechanical". Both have been ruled as fair use.

Re: Thomson Reuters wins first major AI copyright case in the US

#126
post #23

Here's the full decision, which (like most decisions!) is largely written to be legible to non-lawyers: https://storage.courtlistener.com/recap/gov.uscourts.ded.721... The core story seems to be: Westlaw writes and owns headnotes that help lawyers find legal cases about a particular topic. Ross paid people to translate those headnotes into new text, trained an AI on the translations, and used those to make a model th…

This is an interesting opinion, but there are aspects of it that I doubt will stand the test of time. One aspect is the court’s ruling that West’s headnotes are copyrightable even when they merely quote a court opinion verbatim, because the editorial decision to quote the material itself shows a “creative spark”. It really isn’t workable — in law specifically - for copyright to attach to the mere selection of a quote…

" court’s ruling that West’s headnotes are copyrightable even when they merely quote a court opinion verbatim"

That is the opposite of the ruling. The judge said the ones that summarize and pick out the important parts are copyrightable and specifically excludes the headnotes that quote court opinion verbatim.

The judge:

"But I am still not granting summary judgment on any headnotes that are verbatim copies of the case opinion (for reasons that I explain below)"

Re: Thomson Reuters wins first major AI copyright case in the US

#127
post #8

Earlier quoted context omitted.

rotate similar [but different] fonts [or character pages] over each character. the sequence represents data thus watermark.

but the font changes won't be expressed in the (plain text) output of the LLM.

yes thats right a plain text will be distinctive from a watermark version thus outed as an automated forgery. vs incorrect watermark sugesting human attempt to forge, this introduces complications for the generation of output, namely conserving the cypher as well as making sense

Re: Thomson Reuters wins first major AI copyright case in the US

#128

Earlier quoted context omitted.

This is an interesting opinion, but there are aspects of it that I doubt will stand the test of time. One aspect is the court’s ruling that West’s headnotes are copyrightable even when they merely quote a court opinion verbatim, because the editorial decision to quote the material itself shows a “creative spark”. It really isn’t workable — in law specifically - for copyright to attach to the mere selection of a quote…

" court’s ruling that West’s headnotes are copyrightable even when they merely quote a court opinion verbatim" That is the opposite of the ruling. The judge said the ones that summarize and pick out the important parts are copyrightable and specifically excludes the headnotes that quote court opinion verbatim. The judge: "But I am still not granting summary judgment on any headnotes that are verbatim copies of the ca…

You're right as far as the MSJ is concerned, and I should've been more precise. I was focusing on the dictum in the preceding paragraph (because we're discussing the broader implications of the order rather than the nuts-and-bolts of the instant motion). In that paragraph, the judge wrote:

> More than that, each headnote is an individual, copyrightable work. That became clear to me once I analogized the lawyer’s editorial judgment to that of a sculptor. A block of raw marble, like a judicial opinion, is not copyrightable. Yet a sculptor creates a sculpture by choosing what to cut away and what to leave in place. That sculpture is copyrightable. 17 U.S.C. §102(a)(5). So too, even a headnote taken verbatim from an opinion is a carefully chosen fraction of the whole. Identifying which words matter and chiseling away the surrounding mass expresses the editor’s idea about what the important point of law from the opinion is. That editorial expression has enough “creative spark” to be original. ... So all headnotes, even any that quote judicial opinions verbatim, have original value as individual works.

I personally don't think this sculpture metaphor works for verbatim quotes from judicial opinions.

Re: Thomson Reuters wins first major AI copyright case in the US

#130
post #57

Earlier quoted context omitted.

Interestingly, almost the entirety of the judge's opinion seems to be focused on the question of whether the translated notes are subject to copyright. It seems to completely ignore the question of whether training an AI on copyrighted material constitutes making a copy of that work in the first place. Am I missing something? The judge does note that no copyrighted material was distributed to users, because the AI do…

Ross evidently copied and used the text himself. It's like Ross creating an unauthorized volume of West's books, perhaps with a twist. Obscurity ≠ legal compliance.

So the use of AI actually has nothing to do with the ruling here? This is just about the fact that Ross made one local copy of the notes and never distributed it?
Post reply on HN