Live data from Hacker News

Judge rejects most ChatGPT copyright claims from book authors

arstechnica.com

21–30 of 126 posts

Re: Judge rejects most ChatGPT copyright claims from book authors

#21
Though I think that training data is totally a violation of copyright, OpenAI really needs to win. Copyright has been unreasonably extended to the point where it's untenable and if we needed to track down rights holders and negotiate 'training rights' we'd ensure there would be no open models or competition going forward. The rights holders are going to lose this one and it's probably for the better.

Re: Judge rejects most ChatGPT copyright claims from book authors

#22
post #19

If I memorize a book, am I guilty of copyright infringement?

Yes if you write it back out again.

What if I merely demonstrate knowledge that means I've memorized the book? For example, being able to answer yes/no questions about what's on a particular page?

The point I'm trying to make here is that in this situation I have a representation of the entire book in my brain. Haven't I copied it into my neurons?

Re: Judge rejects most ChatGPT copyright claims from book authors

#23

Though I think that training data is totally a violation of copyright, OpenAI really needs to win. Copyright has been unreasonably extended to the point where it's untenable and if we needed to track down rights holders and negotiate 'training rights' we'd ensure there would be no open models or competition going forward. The rights holders are going to lose this one and it's probably for the better.

> Though I think that training data is totally a violation of copyright, OpenAI really needs to win.

I’d argue that that’s not really how the law is supposed to work.

Re: Judge rejects most ChatGPT copyright claims from book authors

#24

Though I think that training data is totally a violation of copyright, OpenAI really needs to win. Copyright has been unreasonably extended to the point where it's untenable and if we needed to track down rights holders and negotiate 'training rights' we'd ensure there would be no open models or competition going forward. The rights holders are going to lose this one and it's probably for the better.

> Though I think that training data is totally a violation of copyright, OpenAI really needs to win. I’d argue that that’s not really how the law is supposed to work.

Ultimately it's gonna come down to new legislation, and there is just absolutely no chance of "big tech gets to exploit every copyrighted work on the planet for free". Definitely not in the US or Europe.

Big media companies don't want that, zillions of small creators don't want that, the public at large has no great affection for giant tech companies getting even more power.

Re: Judge rejects most ChatGPT copyright claims from book authors

#25

Earlier quoted context omitted.

Is it just me or is that really the main and most consequential claim? Title feels misleading.

Yeah agreed, and also the general direction this is heading in feels a bit strange to me - we can debate the merits of current copyright law - but the indisputable fact is that the training material came from *pirated* copies of these author’s copyrighted works. If we are to believe the much used arguments against piracy that we’ve been fed over the last 30+ years (mainly by deep pocketed media companies) then piracy…

> Otherwise WTF have we been prosecuting individuals for all these years for pirating movies and music etc?

We don't. Nobody got prosecuted for downloading movies. Because that isn't illegal.

Whats illegal is distributing copies to other people.

Its OK though. Its a common misconception that internet piracy is illegal.

Only the distribution part has ever been successfully prosecuted.

> They didn’t do that. They used stolen copies instead.

Using stolen copies isn't illegal. It is completely legal to use other people's work without their permission. Distributing copies is the only part that is illegal.

> if we just throw copyright law completely out of the window at the same time.

Nothing is being throw out the window. The facts that I have described have always been true.

Re: Judge rejects most ChatGPT copyright claims from book authors

#26
While I'm not optimistic about AI in general, if they're going to train LLMs I think they ought to use good data, and published books may on average be better for that than scraping the web.

Relatedly, I wonder how many humans have ever learned anything from pirated ebooks --- in some countries, I bet that number is close to 100%.

Re: Judge rejects most ChatGPT copyright claims from book authors

#27
post #22

Earlier quoted context omitted.

Yes if you write it back out again.

What if I merely demonstrate knowledge that means I've memorized the book? For example, being able to answer yes/no questions about what's on a particular page? The point I'm trying to make here is that in this situation I have a representation of the entire book in my brain. Haven't I copied it into my neurons?

When ChatGPT gets agency regarding what goes into training data and what doesn't, gets paid minimum wage for the work it does, and can be held liable for violations of the law, your question will be very relevant.

Until then...

Re: Judge rejects most ChatGPT copyright claims from book authors

#29

It's completely unsurprising that the entire train of extremely harsh, heavy handed and heavily-applied arguments about using pirated works and copyrighted materials in any way without their holder's permission get so commonly applied to ordinary people and small players of any kind, only to be swept away in legalese when the creators themselves try to use them against a large well connected industry with deep pocket…

Totally agree. Though if they follow this to the end then there will at least be a legal precedent to point to for all future piracy prosecutions.

“As shown in Silverman v OpenAI [2024], copyright is no longer enforceable….”

Re: Judge rejects most ChatGPT copyright claims from book authors

#30

Though I think that training data is totally a violation of copyright, OpenAI really needs to win. Copyright has been unreasonably extended to the point where it's untenable and if we needed to track down rights holders and negotiate 'training rights' we'd ensure there would be no open models or competition going forward. The rights holders are going to lose this one and it's probably for the better.

Why do you think training data is a violation of copyright? What is being copied?
Post reply on HN