Live data from Hacker News

Judge rejects most ChatGPT copyright claims from book authors

arstechnica.com

41–50 of 126 posts

Re: Judge rejects most ChatGPT copyright claims from book authors

#41

Though I think that training data is totally a violation of copyright, OpenAI really needs to win. Copyright has been unreasonably extended to the point where it's untenable and if we needed to track down rights holders and negotiate 'training rights' we'd ensure there would be no open models or competition going forward. The rights holders are going to lose this one and it's probably for the better.

Why do you think training data is a violation of copyright? What is being copied?

Did OpenAI buy a copy of every book in The Pile? Or did they all just fall off the back of a truck?

Aside from that question, I tend to agree with the judge - LLM outputs are obviously not derivative works or copies in any sense that we normally use the phrase.

Re: Judge rejects most ChatGPT copyright claims from book authors

#42

Earlier quoted context omitted.

Yeah agreed, and also the general direction this is heading in feels a bit strange to me - we can debate the merits of current copyright law - but the indisputable fact is that the training material came from *pirated* copies of these author’s copyrighted works. If we are to believe the much used arguments against piracy that we’ve been fed over the last 30+ years (mainly by deep pocketed media companies) then piracy…

> Otherwise WTF have we been prosecuting individuals for all these years for pirating movies and music etc? We don't. Nobody got prosecuted for downloading movies. Because that isn't illegal. Whats illegal is distributing copies to other people. Its OK though. Its a common misconception that internet piracy is illegal. Only the distribution part has ever been successfully prosecuted. > They didn’t do that. They used…

It’s not just distribution. Making unauthorised reproductions of copyrighted work is also copyright infringement.

I take your point that prosecutions are focused more on the sharers than individual downloaders, but OpenAI still infringed copyright by making unauthorised reproductions of the authors’ works (by downloading pirated copies of them).

Re: Judge rejects most ChatGPT copyright claims from book authors

#43

Though I think that training data is totally a violation of copyright, OpenAI really needs to win. Copyright has been unreasonably extended to the point where it's untenable and if we needed to track down rights holders and negotiate 'training rights' we'd ensure there would be no open models or competition going forward. The rights holders are going to lose this one and it's probably for the better.

"training data is totally a violation of copyright"

This really isn't clear because cognition is treated as a special exception to copyright. Every thought we have is derivative of everything we've seen before to some degree; reading a book makes our brains a derivative work. But we recognize that cognition is special.

With machines we tend to apply a strict test: Did copyright go in? If so, the output is almost certainly derivative.

With human brains, with cognition, it isn't enough to prove that a person has consumed a copywitten work prior to having a thought -- instead we judge every thought individually as to its originality.

If we are in a position to apply similar cognitive rules to an LLM then the weights won't be derivative works and we will judge each output as to its originality rather than simply assume.

Re: Judge rejects most ChatGPT copyright claims from book authors

#44
post #16
post #9

Earlier quoted context omitted.

The piracy argument can be fixed by OpenAI buying one copy of each work. The overall question of whether they're allowed to train on copyrighted material without permission seems much larger and more interesting.

If learning from a purchased, copyrighted work is illegal, colleges are in real trouble. Textbook publishers will be thrilled though: this book is $200 to read, but you need an additional license to learn anything from it.

Learning and then reproducing parts of a work already is illegal, depending on context.

Also, “learning” here isn’t the same thing as what college students do. For one thing, you can’t copy-paste a college student’s whole brain. You can’t own and sell their brains (well, ah, you know what I mean—not the organ trade). For another… it’s simply not the same thing. It might be, we suppose, similar to some parts of how human learning work, but it’s plainly not identical and may not be especially close. We use the same word for what LLMs do because it’s a convenient and useful-enough analogy, but that doesn’t mean we can prove anything else simply based on our having re-used the word “learn” here.

“That’s learning, this is learning, so all things that apply to one apply to the other”—no, that doesn’t follow, it takes more than that.

Re: Judge rejects most ChatGPT copyright claims from book authors

#45
post #38

Though I think that training data is totally a violation of copyright, OpenAI really needs to win. Copyright has been unreasonably extended to the point where it's untenable and if we needed to track down rights holders and negotiate 'training rights' we'd ensure there would be no open models or competition going forward. The rights holders are going to lose this one and it's probably for the better.

> Though I think that training data is totally a violation of copyright, OpenAI really needs to win. It will have 2 consequences: 1. No reliable law. 2. No win for humans. Only big corps (openai and corporations behind it) wins. Is it really what we all deserve?

And what do you think the consequence will be if Open Ai loses?

A: Big media companies will make billions from licensing, AI will only be produced by a few companies that can afford the licenses and creators will get an annual $10 check (see Spotify).

Re: Judge rejects most ChatGPT copyright claims from book authors

#46

Earlier quoted context omitted.

Is it just me or is that really the main and most consequential claim? Title feels misleading.

Yeah agreed, and also the general direction this is heading in feels a bit strange to me - we can debate the merits of current copyright law - but the indisputable fact is that the training material came from *pirated* copies of these author’s copyrighted works. If we are to believe the much used arguments against piracy that we’ve been fed over the last 30+ years (mainly by deep pocketed media companies) then piracy…

Fair use is non-transitive: you reviewing a pirated copy of a movie can be fair use even if that copy isn't. If training on copyrighted images is fair use then it doesn't matter how you got those images. Think about it this way: if the opposite were true, then being able to review a movie would be a privilege you have to pay for by buying the movie, rather than just something you can do because the 1st Amendment exists.

A concrete example of this is Google Images. Image search is fair use, even though they index shittons of infringing images.

The judge isn't going to touch the "output is infringing" argument mainly because the authors didn't actually connect the dots between their work and a specific output. That's an argument that would stick if you were suing a user of OpenAI's services, not OpenAI themselves. "Output is infringing" is going to be part of the fair use analysis anyway (specifically, the market substitution factor), so this extra claim is superfluous.

Re: Judge rejects most ChatGPT copyright claims from book authors

#47

Though I think that training data is totally a violation of copyright, OpenAI really needs to win. Copyright has been unreasonably extended to the point where it's untenable and if we needed to track down rights holders and negotiate 'training rights' we'd ensure there would be no open models or competition going forward. The rights holders are going to lose this one and it's probably for the better.

"training data is totally a violation of copyright" This really isn't clear because cognition is treated as a special exception to copyright. Every thought we have is derivative of everything we've seen before to some degree; reading a book makes our brains a derivative work. But we recognize that cognition is special. With machines we tend to apply a strict test: Did copyright go in? If so, the output is almost cert…

"This really isn't clear because cognition is treated as a special exception to copyright."

Actually, no. It's considered a transformative use. If you memorize a copyrighted play or piece of music and then perform in in public, that's a copyright violation. It's the literalness of the copy that matters.

Re: Judge rejects most ChatGPT copyright claims from book authors

#48

Though I think that training data is totally a violation of copyright, OpenAI really needs to win. Copyright has been unreasonably extended to the point where it's untenable and if we needed to track down rights holders and negotiate 'training rights' we'd ensure there would be no open models or competition going forward. The rights holders are going to lose this one and it's probably for the better.

"training data is totally a violation of copyright" This really isn't clear because cognition is treated as a special exception to copyright. Every thought we have is derivative of everything we've seen before to some degree; reading a book makes our brains a derivative work. But we recognize that cognition is special. With machines we tend to apply a strict test: Did copyright go in? If so, the output is almost cert…

>> "training data is totally a violation of copyright"

> This really isn't clear because cognition is treated as a special exception to copyright.

Human cognition; not the latest algorithms and their output, which some enthusiastic software engineers eagerly confuse for cognition. It's actually pretty clear.

Re: Judge rejects most ChatGPT copyright claims from book authors

#49

Though I think that training data is totally a violation of copyright, OpenAI really needs to win. Copyright has been unreasonably extended to the point where it's untenable and if we needed to track down rights holders and negotiate 'training rights' we'd ensure there would be no open models or competition going forward. The rights holders are going to lose this one and it's probably for the better.

> Though I think that training data is totally a violation of copyright, OpenAI really needs to win. I’d argue that that’s not really how the law is supposed to work.

Well, the objection to current copyright laws is that those laws are not how copyright ought to work.

Re: Judge rejects most ChatGPT copyright claims from book authors

#50
post #11

I have a more general question. Say I read a science fiction book which has descriptions of some futuristic technologies. I get inspired by it and spend lot of my time and energy becoming an expert in the required engineering and technologies and invent a machine/process to make that futuristic technology a reality. If I attempt to commercialize my work, can I get sued for copyright infringement by the author(s) of t…

Several science fiction authors wrote about space rockets and satellites well before anyone actually built them in the 1950s.
Post reply on HN