Live data from Hacker News

Judge rejects most ChatGPT copyright claims from book authors

arstechnica.com

61–70 of 126 posts

Re: Judge rejects most ChatGPT copyright claims from book authors

#61

Earlier quoted context omitted.

Yeah agreed, and also the general direction this is heading in feels a bit strange to me - we can debate the merits of current copyright law - but the indisputable fact is that the training material came from *pirated* copies of these author’s copyrighted works. If we are to believe the much used arguments against piracy that we’ve been fed over the last 30+ years (mainly by deep pocketed media companies) then piracy…

Fair use is non-transitive: you reviewing a pirated copy of a movie can be fair use even if that copy isn't. If training on copyrighted images is fair use then it doesn't matter how you got those images. Think about it this way: if the opposite were true, then being able to review a movie would be a privilege you have to pay for by buying the movie, rather than just something you can do because the 1st Amendment exis…

It has not been proven that AI training data is fair use.

Re: Judge rejects most ChatGPT copyright claims from book authors

#62
> Arguing that OpenAI caused economic injury by unfairly repurposing authors' works, even if authors could show evidence of a DMCA violation, authors could only speculate about what injury was caused

That's an interesting angle I hadn't seen coming, but it makes sense: the system isn't reliable enough to output a book proper, so you'd never use it to read a book (can't ask it to output lord of the rings page 47 ... yet), and summaries were always okay to post as far as I know

So if I post excerpts of books that neither I nor the querier has a license for when they guess/engineer a search query, even if that (with lots of effort) can amount to the whole book, that's apparently not causing any economic injury to the author. It's not ruled to not be copyright infringement yet, so maybe you can force removal from the market but you can't be awarded damages as I understand it?

Separately, I'm a bit surprised the authors haven't done much in the way of discovery and just alleged baselsssly that OpenAI's training process removed copyright notices. The judge says it's unsubstantiated and there's even counter-evidence. Wouldn't it be a simple matter to subpoena a list of works/sources that were used and under which license? That would also be interesting to learn for others working in the field

Re: Judge rejects most ChatGPT copyright claims from book authors

#63
post #22

Earlier quoted context omitted.

Yes if you write it back out again.

What if I merely demonstrate knowledge that means I've memorized the book? For example, being able to answer yes/no questions about what's on a particular page? The point I'm trying to make here is that in this situation I have a representation of the entire book in my brain. Haven't I copied it into my neurons?

If we're technical, there is a substantial legal difference between your brain and an AI system, because the criteria for what counts as a copy is defined (in US) as "“Copies” are material objects, other than phonorecords, in which a work is fixed by any method now known or later developed, and from which the work can be perceived, reproduced, or otherwise communicated, either directly or with the aid of a machine or device. The term “copies” includes the material object, other than a phonorecord, in which the work is first fixed." and all the law precedent doesn't consider any representation of something in your brain as "the work being fixed" , so that isn't a copy and anything that applies to copies doesn't apply to your memory, but do apply to any computer memory or any future technical method to make that representation.

So analogies to the brain aren't appropriate, because human memories are special and separate in the eyes of law, and if we'd make a machine that does literally exactly the same thing as your brain does, it will still NOT have the same legal treatment as your brain; there is no legal principle that it should get equal or similar treatment.

Re: Judge rejects most ChatGPT copyright claims from book authors

#64
post #22

Earlier quoted context omitted.

Yes if you write it back out again.

What if I merely demonstrate knowledge that means I've memorized the book? For example, being able to answer yes/no questions about what's on a particular page? The point I'm trying to make here is that in this situation I have a representation of the entire book in my brain. Haven't I copied it into my neurons?

Do you profit from people testing your ability to memorise the book and give them the ability to substantially recreate the original work and act in competition to the original author?

Re: Judge rejects most ChatGPT copyright claims from book authors

#65
post #47

Earlier quoted context omitted.

"training data is totally a violation of copyright" This really isn't clear because cognition is treated as a special exception to copyright. Every thought we have is derivative of everything we've seen before to some degree; reading a book makes our brains a derivative work. But we recognize that cognition is special. With machines we tend to apply a strict test: Did copyright go in? If so, the output is almost cert…

"This really isn't clear because cognition is treated as a special exception to copyright." Actually, no. It's considered a transformative use. If you memorize a copyrighted play or piece of music and then perform in in public, that's a copyright violation. It's the literalness of the copy that matters.

No, that's totally incorrect, we do not consider every observation a "transformative use" as applied to the human mind. If you memorize a copyrighted play and write another play it is NOT inherently a copyright violation of everything which has come before. We just don't do that.

The new play is judged as to its originality.

People who have seen a play (everybody) are allowed to write new plays which aren't beholden to the copyright of the first play they've ever watched.

Re: Judge rejects most ChatGPT copyright claims from book authors

#66

Earlier quoted context omitted.

"training data is totally a violation of copyright" This really isn't clear because cognition is treated as a special exception to copyright. Every thought we have is derivative of everything we've seen before to some degree; reading a book makes our brains a derivative work. But we recognize that cognition is special. With machines we tend to apply a strict test: Did copyright go in? If so, the output is almost cert…

1) no similarities have ever been demonstrated between large language models and human cognition, and until that happens (spoiler: never) there is no basis in comparing them like this. 2) even if they were somehow proven to be the same there is still no reason why the same standards need to be applied to computer programs and humans because computer programs do not have any rights or legal protections. 3) cognition i…

"1) no similarities have ever been demonstrated between large language models and human cognition"

This is false. The LLM's entire purpose is to mimic cognition.

You could argue that the operation differs in important ways - of course. But the similarity of output is literally the entire point.

"2) even if they were somehow proven to be the same"

I didn't suggest they need to be the same, proven or otherwise. I think you're not understanding. The point is that the function is similar.

How it works doesn't necessarily matter.

"3) cognition is not a "special exception to copyright" because it is entirely unrelated. "

False as a matter of law.

"4) we do not "judge every thought individually as to it's originality" because other peoples' thoughts are entirely opaque."

Also false as a matter of law. When you publish your thoughts - your works, writing, whatever they are judged as to their originality if the question of who owns the copyright is raised.

"Nobody is judging your thoughts, and if you think they are you need to take your medications."

There's no need to be snarky and disingenuous.

From the comment guidelines: Be kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes.

https://news.ycombinator.com/newsguidelines.html

Re: Judge rejects most ChatGPT copyright claims from book authors

#67
post #45

Earlier quoted context omitted.

And what do you think the consequence will be if Open Ai loses? A: Big media companies will make billions from licensing, AI will only be produced by a few companies that can afford the licenses and creators will get an annual $10 check (see Spotify).

You forgot the last point, which is that creators who don't want their work used in the training data for these megacorp LLMs without permission will get what they want.

And we won't have open models.

Re: Judge rejects most ChatGPT copyright claims from book authors

#68
post #48

Earlier quoted context omitted.

"training data is totally a violation of copyright" This really isn't clear because cognition is treated as a special exception to copyright. Every thought we have is derivative of everything we've seen before to some degree; reading a book makes our brains a derivative work. But we recognize that cognition is special. With machines we tend to apply a strict test: Did copyright go in? If so, the output is almost cert…

>> "training data is totally a violation of copyright" > This really isn't clear because cognition is treated as a special exception to copyright. Human cognition; not the latest algorithms and their output, which some enthusiastic software engineers eagerly confuse for cognition. It's actually pretty clear.

Carbon chauvinism at its finest.

Re: Judge rejects most ChatGPT copyright claims from book authors

#69

Earlier quoted context omitted.

Yeah agreed, and also the general direction this is heading in feels a bit strange to me - we can debate the merits of current copyright law - but the indisputable fact is that the training material came from *pirated* copies of these author’s copyrighted works. If we are to believe the much used arguments against piracy that we’ve been fed over the last 30+ years (mainly by deep pocketed media companies) then piracy…

> Otherwise WTF have we been prosecuting individuals for all these years for pirating movies and music etc? We don't. Nobody got prosecuted for downloading movies. Because that isn't illegal. Whats illegal is distributing copies to other people. Its OK though. Its a common misconception that internet piracy is illegal. Only the distribution part has ever been successfully prosecuted. > They didn’t do that. They used…

Of course it is illegal in many countries to download pirated movies. And the USA put a lot of pressure on countries where it isn't illegal - I know that well enough, in my country it was and is legal.

Re: Judge rejects most ChatGPT copyright claims from book authors

#70
post #48

Earlier quoted context omitted.

"training data is totally a violation of copyright" This really isn't clear because cognition is treated as a special exception to copyright. Every thought we have is derivative of everything we've seen before to some degree; reading a book makes our brains a derivative work. But we recognize that cognition is special. With machines we tend to apply a strict test: Did copyright go in? If so, the output is almost cert…

>> "training data is totally a violation of copyright" > This really isn't clear because cognition is treated as a special exception to copyright. Human cognition; not the latest algorithms and their output, which some enthusiastic software engineers eagerly confuse for cognition. It's actually pretty clear.

As I said, human cognition is a special case.

The open question is how to handle machines that mimic the process.

Post reply on HN