Earlier quoted context omitted.
but the font changes won't be expressed in the (plain text) output of the LLM.
Presumably the font will represent letters to look like a different letter, making it not useful to LLMs scraping the site but useful for visual readers. This would have detrimental effects to people who use screen readers or have their own stylesheets of course.
Thomson Reuters wins first major AI copyright case in the US
121–130 of 188 posts
Re: Thomson Reuters wins first major AI copyright case in the US
#122See. The fair-use excuses that the AI proponents here were trying to hang on to for dear life have fallen flat on this ruling. This is going to be one of many cases in which there will be licensing deals being made out of this to stop AI grifters claiming 'fair use' to try to side-step copyright laws because they are using a gen AI system. OpenAI ended up paying up for the data with Shutterstock and other news source…
My biggest concern is, what happens when countries like China, who aren't restricted by this, far outpace western countries in this technology? Do we just shrug and accept our far inferior models? LLMs are a productivity multiplier (similar to a search engine), so it'll have a large impact on the economy if licensing costs prohibit large scale training.
Re: Thomson Reuters wins first major AI copyright case in the US
#123Thomson Reuters chose to sue Ross Intelligence, not a company like Google or even OpenAI. I wonder how deeper pockets would have affected the outcome. I wonder how the politics played out. The big AI companies could have funded Ross Intelligence, who could have threatened to sabotage their legal strategies by tanking and settling their own case in TR's favor.
Re: Thomson Reuters wins first major AI copyright case in the US
#124Earlier quoted context omitted.
> So the models are legitimately not viable without massive copyright infringement. Copyright is not about acquisition, it is about publication and/or distribution. If I get a copy of Harry Potter from a dumpster, I can read it. If a company gets a copy of *all books from a torrent, they can use it to train their AI. The torrent providers may be in violation of copyright, and if the AI can be used to reproduce substa…
That's an interesting take, but false in a lot of juristictions. Even if we ignore question of if the model can distribute work, in many places even downloading content is illegal. Otherwise the person torrenting a movie would be totally in the clear, or thing about what MS would say if a company "just" downloads copies of Windows to use on their computers without ever distributing them.
Any examples of people being sued for merely downloading? "Torrenting" basically always involves uploading, even if you stop immediately after completion. A better test would be if someone was sued for using an illegal streaming site, which to my knowledge has never happened.
Re: Thomson Reuters wins first major AI copyright case in the US
#125Earlier quoted context omitted.
> So the models are legitimately not viable without massive copyright infringement. Copyright is not about acquisition, it is about publication and/or distribution. If I get a copy of Harry Potter from a dumpster, I can read it. If a company gets a copy of *all books from a torrent, they can use it to train their AI. The torrent providers may be in violation of copyright, and if the AI can be used to reproduce substa…
> simply training a model on illegally distributed text should not be copyright infringement You can train a model on copyrighted text, you just can't distribute the output in any way without violating copyright. (edit: depending on the other fair use factors). One of the big problems is that training is a mechanical process, so there is a direct line between the copyrighted works and the model's output, regardless o…
Google making thumbnails or scanning books are both arguably "mechanical". Both have been ruled as fair use.
Re: Thomson Reuters wins first major AI copyright case in the US
#126Here's the full decision, which (like most decisions!) is largely written to be legible to non-lawyers: https://storage.courtlistener.com/recap/gov.uscourts.ded.721... The core story seems to be: Westlaw writes and owns headnotes that help lawyers find legal cases about a particular topic. Ross paid people to translate those headnotes into new text, trained an AI on the translations, and used those to make a model th…
This is an interesting opinion, but there are aspects of it that I doubt will stand the test of time. One aspect is the court’s ruling that West’s headnotes are copyrightable even when they merely quote a court opinion verbatim, because the editorial decision to quote the material itself shows a “creative spark”. It really isn’t workable — in law specifically - for copyright to attach to the mere selection of a quote…
That is the opposite of the ruling. The judge said the ones that summarize and pick out the important parts are copyrightable and specifically excludes the headnotes that quote court opinion verbatim.
The judge:
"But I am still not granting summary judgment on any headnotes that are verbatim copies of the case opinion (for reasons that I explain below)"
Re: Thomson Reuters wins first major AI copyright case in the US
#127Earlier quoted context omitted.
rotate similar [but different] fonts [or character pages] over each character. the sequence represents data thus watermark.
but the font changes won't be expressed in the (plain text) output of the LLM.
Re: Thomson Reuters wins first major AI copyright case in the US
#128Earlier quoted context omitted.
This is an interesting opinion, but there are aspects of it that I doubt will stand the test of time. One aspect is the court’s ruling that West’s headnotes are copyrightable even when they merely quote a court opinion verbatim, because the editorial decision to quote the material itself shows a “creative spark”. It really isn’t workable — in law specifically - for copyright to attach to the mere selection of a quote…
" court’s ruling that West’s headnotes are copyrightable even when they merely quote a court opinion verbatim" That is the opposite of the ruling. The judge said the ones that summarize and pick out the important parts are copyrightable and specifically excludes the headnotes that quote court opinion verbatim. The judge: "But I am still not granting summary judgment on any headnotes that are verbatim copies of the ca…
> More than that, each headnote is an individual, copyrightable work. That became clear to me once I analogized the lawyer’s editorial judgment to that of a sculptor. A block of raw marble, like a judicial opinion, is not copyrightable. Yet a sculptor creates a sculpture by choosing what to cut away and what to leave in place. That sculpture is copyrightable. 17 U.S.C. §102(a)(5). So too, even a headnote taken verbatim from an opinion is a carefully chosen fraction of the whole. Identifying which words matter and chiseling away the surrounding mass expresses the editor’s idea about what the important point of law from the opinion is. That editorial expression has enough “creative spark” to be original. ... So all headnotes, even any that quote judicial opinions verbatim, have original value as individual works.
I personally don't think this sculpture metaphor works for verbatim quotes from judicial opinions.
Re: Thomson Reuters wins first major AI copyright case in the US
#129Re: Thomson Reuters wins first major AI copyright case in the US
#130Earlier quoted context omitted.
Interestingly, almost the entirety of the judge's opinion seems to be focused on the question of whether the translated notes are subject to copyright. It seems to completely ignore the question of whether training an AI on copyrighted material constitutes making a copy of that work in the first place. Am I missing something? The judge does note that no copyrighted material was distributed to users, because the AI do…
Ross evidently copied and used the text himself. It's like Ross creating an unauthorized volume of West's books, perhaps with a twist. Obscurity ≠ legal compliance.