Live data from Hacker News

Judge rejects most ChatGPT copyright claims from book authors

arstechnica.com

51–60 of 126 posts

Re: Judge rejects most ChatGPT copyright claims from book authors

#51

Though I think that training data is totally a violation of copyright, OpenAI really needs to win. Copyright has been unreasonably extended to the point where it's untenable and if we needed to track down rights holders and negotiate 'training rights' we'd ensure there would be no open models or competition going forward. The rights holders are going to lose this one and it's probably for the better.

"training data is totally a violation of copyright" This really isn't clear because cognition is treated as a special exception to copyright. Every thought we have is derivative of everything we've seen before to some degree; reading a book makes our brains a derivative work. But we recognize that cognition is special. With machines we tend to apply a strict test: Did copyright go in? If so, the output is almost cert…

1) no similarities have ever been demonstrated between large language models and human cognition, and until that happens (spoiler: never) there is no basis in comparing them like this.

2) even if they were somehow proven to be the same there is still no reason why the same standards need to be applied to computer programs and humans because computer programs do not have any rights or legal protections.

3) cognition is not a "special exception to copyright" because it is entirely unrelated. "Copy" "right" is who has rights to make copies. Your thoughts are not considered copies because they are intangible.

4) we do not "judge every thought individually as to it's originality" because other peoples' thoughts are entirely opaque. Nobody is judging your thoughts, and if you think they are you need to take your medications.

Re: Judge rejects most ChatGPT copyright claims from book authors

#52
post #11

I have a more general question. Say I read a science fiction book which has descriptions of some futuristic technologies. I get inspired by it and spend lot of my time and energy becoming an expert in the required engineering and technologies and invent a machine/process to make that futuristic technology a reality. If I attempt to commercialize my work, can I get sued for copyright infringement by the author(s) of t…

I believe the actual law is the clearest answer to your question:

US copyright law, Subject matter of copyright 102.(b) "In no case does copyright protection for an original work of authorship extend to any idea, procedure, process, system, method of operation, concept, principle, or discovery, regardless of the form in which it is described, explained, illustrated, or embodied in such work." (https://www.copyright.gov/title17/92chap1.html)

Quite explicitly, copyright law doesn't grant the author any exclusive rights to the ideas expressed in the work, just on the particular creative expression.

Re: Judge rejects most ChatGPT copyright claims from book authors

#53

Earlier quoted context omitted.

"training data is totally a violation of copyright" This really isn't clear because cognition is treated as a special exception to copyright. Every thought we have is derivative of everything we've seen before to some degree; reading a book makes our brains a derivative work. But we recognize that cognition is special. With machines we tend to apply a strict test: Did copyright go in? If so, the output is almost cert…

1) no similarities have ever been demonstrated between large language models and human cognition, and until that happens (spoiler: never) there is no basis in comparing them like this. 2) even if they were somehow proven to be the same there is still no reason why the same standards need to be applied to computer programs and humans because computer programs do not have any rights or legal protections. 3) cognition i…

Thank you for saying what I was going to say to this person. I'm so fucking tired of seeing people who probably have never opened a neuroscience textbook talk about cognition.

Re: Judge rejects most ChatGPT copyright claims from book authors

#54
post #45
post #38

Earlier quoted context omitted.

> Though I think that training data is totally a violation of copyright, OpenAI really needs to win. It will have 2 consequences: 1. No reliable law. 2. No win for humans. Only big corps (openai and corporations behind it) wins. Is it really what we all deserve?

And what do you think the consequence will be if Open Ai loses? A: Big media companies will make billions from licensing, AI will only be produced by a few companies that can afford the licenses and creators will get an annual $10 check (see Spotify).

You forgot the last point, which is that creators who don't want their work used in the training data for these megacorp LLMs without permission will get what they want.

Re: Judge rejects most ChatGPT copyright claims from book authors

#55

Though I think that training data is totally a violation of copyright, OpenAI really needs to win. Copyright has been unreasonably extended to the point where it's untenable and if we needed to track down rights holders and negotiate 'training rights' we'd ensure there would be no open models or competition going forward. The rights holders are going to lose this one and it's probably for the better.

Why do you think training data is a violation of copyright? What is being copied?

OpenAI obviously copied the data onto their servers when they collected the training data.

Re: Judge rejects most ChatGPT copyright claims from book authors

#57
post #45
post #38

Earlier quoted context omitted.

> Though I think that training data is totally a violation of copyright, OpenAI really needs to win. It will have 2 consequences: 1. No reliable law. 2. No win for humans. Only big corps (openai and corporations behind it) wins. Is it really what we all deserve?

And what do you think the consequence will be if Open Ai loses? A: Big media companies will make billions from licensing, AI will only be produced by a few companies that can afford the licenses and creators will get an annual $10 check (see Spotify).

Let my big tech company fuck you or the big media company will fuck you instead.

Re: Judge rejects most ChatGPT copyright claims from book authors

#58

Earlier quoted context omitted.

Why do you think training data is a violation of copyright? What is being copied?

Did OpenAI buy a copy of every book in The Pile? Or did they all just fall off the back of a truck? Aside from that question, I tend to agree with the judge - LLM outputs are obviously not derivative works or copies in any sense that we normally use the phrase.

OpenAI hasn’t trained on The Pile, as far as I know. I think you mean "Did Meta buy a copy of every book in books3?" since llama wasn’t trained on The Pile either. And the answer is no.

It seems important for the answer to remain no, otherwise the only entities that can afford to train on sufficient numbers of books will be big companies. No one can afford 190,000 books except huge corporations, so we’ll be surrendering our ability to train our own useful models if the price of training is that high.

190k books isn’t even that much. OpenAI probably trained on millions.

Re: Judge rejects most ChatGPT copyright claims from book authors

#59

Earlier quoted context omitted.

Yeah agreed, and also the general direction this is heading in feels a bit strange to me - we can debate the merits of current copyright law - but the indisputable fact is that the training material came from *pirated* copies of these author’s copyrighted works. If we are to believe the much used arguments against piracy that we’ve been fed over the last 30+ years (mainly by deep pocketed media companies) then piracy…

> Otherwise WTF have we been prosecuting individuals for all these years for pirating movies and music etc? We don't. Nobody got prosecuted for downloading movies. Because that isn't illegal. Whats illegal is distributing copies to other people. Its OK though. Its a common misconception that internet piracy is illegal. Only the distribution part has ever been successfully prosecuted. > They didn’t do that. They used…

You're conflating different matters here by merging together "legal/illegal" and "successfully prosecuted" - in general, there is a big gap of things prohibited by law which can't and won't ever be prosecuted by the state as they aren't felonies or misdemeanors, but only justify civil claims of compensation, if the other party wants to sue them and cares to put in the effort (and money) to actually do so. I think it's not appropriate to call all the latter scenarios "legal".

Re: Judge rejects most ChatGPT copyright claims from book authors

#60

While I'm not optimistic about AI in general, if they're going to train LLMs I think they ought to use good data, and published books may on average be better for that than scraping the web. Relatedly, I wonder how many humans have ever learned anything from pirated ebooks --- in some countries, I bet that number is close to 100%.

How humans do anything is immaterial to this question.
Post reply on HN