Live data from Hacker News

Judge rejects most ChatGPT copyright claims from book authors

arstechnica.com

71–80 of 126 posts

Re: Judge rejects most ChatGPT copyright claims from book authors

#71
post #62

> Arguing that OpenAI caused economic injury by unfairly repurposing authors' works, even if authors could show evidence of a DMCA violation, authors could only speculate about what injury was caused That's an interesting angle I hadn't seen coming, but it makes sense: the system isn't reliable enough to output a book proper, so you'd never use it to read a book (can't ask it to output lord of the rings page 47 ... y…

> Arguing that OpenAI caused economic injury by unfairly repurposing authors' works, even if authors could show evidence of a DMCA violation, authors could only speculate about what injury was caused

Funny that inability to prove speculated injuries didn't stop judges from ordering people to pay hundreds of thousands in damages in piracy cases

Re: Judge rejects most ChatGPT copyright claims from book authors

#72
post #45
post #38

Earlier quoted context omitted.

> Though I think that training data is totally a violation of copyright, OpenAI really needs to win. It will have 2 consequences: 1. No reliable law. 2. No win for humans. Only big corps (openai and corporations behind it) wins. Is it really what we all deserve?

And what do you think the consequence will be if Open Ai loses? A: Big media companies will make billions from licensing, AI will only be produced by a few companies that can afford the licenses and creators will get an annual $10 check (see Spotify).

B: Development of AI accelerates. Necessity being mother of invention and all that. After all a human doesn't need to ingest every written word out there to show signs of intelligence.

Re: Judge rejects most ChatGPT copyright claims from book authors

#73
post #56

How is this different from the authors of (business book) seeking royalties from (bigtech CEO) because he/she ingested knowledge (training) from this book? Humankind needs to build on top of each other’s knowledge to keep evolving.

Well the first difference is that copyright doesn't protect ideas, only specific texts and their reproduction.

Re: Judge rejects most ChatGPT copyright claims from book authors

#74

Earlier quoted context omitted.

Human brains are not large language models.

Large language models are also not databases of text.

Thought question, not entirely related but if you want to go that that route it actually is.

If I generate some media in say Photoshop. I then send you a JPEG representation of said media. You then distribute a PNG copy of the image without license. Have you violated copyright law?

At what point is there enough parameters to an LLM that it is effectively just a compressed version.

How about deduplicated storage? Is an image stored on that and the reproduced using an index of some sort to distribute a violation.

If I put data into a thing,lets call it training and lets call the thing a model, and then I request the data out of it and get what is perceived as an exact replication of said thing did I create a copy.

Does it matter if we call the thing a hard drive instead?

Re: Judge rejects most ChatGPT copyright claims from book authors

#75

Earlier quoted context omitted.

Did OpenAI buy a copy of every book in The Pile? Or did they all just fall off the back of a truck? Aside from that question, I tend to agree with the judge - LLM outputs are obviously not derivative works or copies in any sense that we normally use the phrase.

OpenAI hasn’t trained on The Pile, as far as I know. I think you mean "Did Meta buy a copy of every book in books3?" since llama wasn’t trained on The Pile either. And the answer is no. It seems important for the answer to remain no, otherwise the only entities that can afford to train on sufficient numbers of books will be big companies. No one can afford 190,000 books except huge corporations, so we’ll be surrender…

If the only issue is paying for a copy, thinking out loud, It would be interesting if some ebook lender like libby let you queue up a whole bunch of books to borrow, train on the return it.

Re: Judge rejects most ChatGPT copyright claims from book authors

#76
post #37

Earlier quoted context omitted.

Does that mean you could rewrite, say, Harry Potter from scratch without violating copyright?

Yes, but unless you change names it would be a trademark violation.

Normally that's the case even though at least one of those books was taken down based on (rather dubious imo) violation of rights on derivative works. It was never a verbatim rewrite nor a trademark violation. The Dutch court even proceeded to make a list of similar ideas between two books which I think is explicitly disallowed under the US version of the copyright law.

https://en.m.wikipedia.org/wiki/Tanya_Grotter

Re: Judge rejects most ChatGPT copyright claims from book authors

#78

Earlier quoted context omitted.

1) no similarities have ever been demonstrated between large language models and human cognition, and until that happens (spoiler: never) there is no basis in comparing them like this. 2) even if they were somehow proven to be the same there is still no reason why the same standards need to be applied to computer programs and humans because computer programs do not have any rights or legal protections. 3) cognition i…

"1) no similarities have ever been demonstrated between large language models and human cognition" This is false. The LLM's entire purpose is to mimic cognition. You could argue that the operation differs in important ways - of course. But the similarity of output is literally the entire point. "2) even if they were somehow proven to be the same" I didn't suggest they need to be the same, proven or otherwise. I think…

>This is false. The LLM's entire purpose is to mimic cognition.

Purpose and mechanism are not the same thing. "Similarity of output" does not make it equivalent.

>I didn't suggest they need to be the same, proven or otherwise. I think you're not understanding. The point is that the function is similar.

Sure, go ahead and ignore all but half a sentence and then accuse me of missing the point.

>False as a matter of law.

Show me the court case where somebody was found to have violated copyright law by thinking about something.

>When you publish your thoughts

You don't publish your thoughts. You publish essays, internet comments, articles, videos, etc based on what you are thinking and those are subject to copyright law.

>There's no need to be snarky and disingenuous.

How dare you, i would never disingenuously tell somebody who thinks his thoughts belong to other people to take their psychiatric medications. Of course i did mean that they should be prescribed by a licensed physician and looking back i regret not stating that explicitly.

Re: Judge rejects most ChatGPT copyright claims from book authors

#79

While I'm not optimistic about AI in general, if they're going to train LLMs I think they ought to use good data, and published books may on average be better for that than scraping the web. Relatedly, I wonder how many humans have ever learned anything from pirated ebooks --- in some countries, I bet that number is close to 100%.

and they can bloody well pay for that like everyone else. get a license for the content that makes your machine that makes you money. it's that simple.

Re: Judge rejects most ChatGPT copyright claims from book authors

#80

Earlier quoted context omitted.

> Otherwise WTF have we been prosecuting individuals for all these years for pirating movies and music etc? We don't. Nobody got prosecuted for downloading movies. Because that isn't illegal. Whats illegal is distributing copies to other people. Its OK though. Its a common misconception that internet piracy is illegal. Only the distribution part has ever been successfully prosecuted. > They didn’t do that. They used…

You're conflating different matters here by merging together "legal/illegal" and "successfully prosecuted" - in general, there is a big gap of things prohibited by law which can't and won't ever be prosecuted by the state as they aren't felonies or misdemeanors, but only justify civil claims of compensation, if the other party wants to sue them and cares to put in the effort (and money) to actually do so. I think it'…

> You're conflating different matters here by merging together "legal/illegal" and "successfully prosecuted"

There isn't any difference between these two things.

If something has never and will never be enforced then it is just words on a piece of paper.

The only way that words on a piece of paper turn into something that actually matters is through enforcement. So the point stands.

It doesn't matter what your interpretation of the words on a piece of paper are if no judge or jury has ever agreed with you and will never agree with you, as evidenced by a successful case.

Post reply on HN