Live data from Hacker News

Judge rejects most ChatGPT copyright claims from book authors

arstechnica.com

1–10 of 126 posts

Re: Judge rejects most ChatGPT copyright claims from book authors

#2
The main excerpt to me:

> The only claim under California's unfair competition law that was allowed to proceed alleged that OpenAI used copyrighted works to train ChatGPT without authors' permission. Because the state law broadly defines what's considered "unfair," Martínez-Olguín said that it's possible that OpenAI's use of the training data "may constitute an unfair practice."

Re: Judge rejects most ChatGPT copyright claims from book authors

#6
post #2

The main excerpt to me: > The only claim under California's unfair competition law that was allowed to proceed alleged that OpenAI used copyrighted works to train ChatGPT without authors' permission. Because the state law broadly defines what's considered "unfair," Martínez-Olguín said that it's possible that OpenAI's use of the training data "may constitute an unfair practice."

Is it just me or is that really the main and most consequential claim? Title feels misleading.

Re: Judge rejects most ChatGPT copyright claims from book authors

#7
post #2

The main excerpt to me: > The only claim under California's unfair competition law that was allowed to proceed alleged that OpenAI used copyrighted works to train ChatGPT without authors' permission. Because the state law broadly defines what's considered "unfair," Martínez-Olguín said that it's possible that OpenAI's use of the training data "may constitute an unfair practice."

Is it just me or is that really the main and most consequential claim? Title feels misleading.

Yeah agreed, and also the general direction this is heading in feels a bit strange to me - we can debate the merits of current copyright law - but the indisputable fact is that the training material came from *pirated* copies of these author’s copyrighted works.

If we are to believe the much used arguments against piracy that we’ve been fed over the last 30+ years (mainly by deep pocketed media companies) then piracy is stealing and if you take stolen material and use it to produce a profitable commercial service then surely that’s a clear cut case in favour of the copyright owners?

Otherwise WTF have we been prosecuting individuals for all these years for pirating movies and music etc?

And here’s the injury: the authors did not get paid for their copyrighted work. Whether you agree with the law or not, OpenAI at the very least should have bought a copy of every book they trained on - and that doesn’t even get into the question of whether they even had the right to train on that material even if they had paid for it. They didn’t do that. They used stolen copies instead.

I don’t understand how it could go any other way. I do get the judge’s reluctance to decide that every bit of output of ChatGPT constitutes copyright infringement - that’s much harder to argue.

I’m not passionately against OpenAI - I pay them money - but I am concerned about this “training material” sourcing issue just being swept under the rug.

AI advancements clearly do impact artists, authors and other creators of works… they’ll be even more screwed if we just throw copyright law completely out of the window at the same time.

Re: Judge rejects most ChatGPT copyright claims from book authors

#8
It's completely unsurprising that the entire train of extremely harsh, heavy handed and heavily-applied arguments about using pirated works and copyrighted materials in any way without their holder's permission get so commonly applied to ordinary people and small players of any kind, only to be swept away in legalese when the creators themselves try to use them against a large well connected industry with deep pockets and lots of lobbying weight.

Suddenly those claiming copyright have become unreasonable and legal contortions of all kinds appear from the blue sky to show why there's little basis to their claims. Never mind that much of the learning material for the LLMs of companies such as OpenAI is explained only opaquely and its legal status (pirated or not, fair use or not, etc) is ambigious at best. Considering that these uses of such material involve a huge load of of obviously commercial aims, it's laughable how biased the legal system has been in their favor so far.

I personally don't favor heavy handed copyright laws in any context but if they're going to be applied with so much snctimony by corporations, academia and governments, then at least the application should be even handed, instead of so blatantly loaded with apparent exemptions.

Re: Judge rejects most ChatGPT copyright claims from book authors

#9

Earlier quoted context omitted.

Is it just me or is that really the main and most consequential claim? Title feels misleading.

Yeah agreed, and also the general direction this is heading in feels a bit strange to me - we can debate the merits of current copyright law - but the indisputable fact is that the training material came from *pirated* copies of these author’s copyrighted works. If we are to believe the much used arguments against piracy that we’ve been fed over the last 30+ years (mainly by deep pocketed media companies) then piracy…

The piracy argument can be fixed by OpenAI buying one copy of each work. The overall question of whether they're allowed to train on copyrighted material without permission seems much larger and more interesting.

Re: Judge rejects most ChatGPT copyright claims from book authors

#10
post #9

Earlier quoted context omitted.

Yeah agreed, and also the general direction this is heading in feels a bit strange to me - we can debate the merits of current copyright law - but the indisputable fact is that the training material came from *pirated* copies of these author’s copyrighted works. If we are to believe the much used arguments against piracy that we’ve been fed over the last 30+ years (mainly by deep pocketed media companies) then piracy…

The piracy argument can be fixed by OpenAI buying one copy of each work. The overall question of whether they're allowed to train on copyrighted material without permission seems much larger and more interesting.

May not be that easy. It may be the case we'll find that (in much the way as tech has made popular ironically) authors/IP creators aren't willing to sell "machine training rights" to their work. A concept that could find itself magicked into existence by artists/publishers.
Post reply on HN