Live data from Hacker News

Judge rejects most ChatGPT copyright claims from book authors

arstechnica.com

101–110 of 126 posts

Re: Judge rejects most ChatGPT copyright claims from book authors

#101
post #99

Earlier quoted context omitted.

Legally if they sell it, they no longer own it and can't determine how it is used. If they license it, that's a different story (they can make a claim of how I breached terms of use, still no copyright issue though). They have absolutely no say in how I use the book they sold to me. Only if I reproduce that book, the actual physical reproduced work is of concern to copyright law, nothing else (my use of the book I bo…

All it takes is a EULA page in the book they sell you to remedy that; a practice well trod by the tech industry to establish the licensing vs. Assumed First Sale relationship. I find it naive to simply assume that publishers won't play tit-for-tat with those that are trying to indirectly monetize their output. I had hope that a return to First Sale sanity may win the day, but as of late I have a dread we're briskly w…

You're 100% and raise a good point.It would only work with digital goods though. At some point people seriously have to realise copyright is just a tool with pro's and con's, a balancing act between differing interests. I agree the principle of first sale should apply even to digital goods. There's some hope in the EU that this is on occasion the case, but much room for improvement everywhere.

Re: Judge rejects most ChatGPT copyright claims from book authors

#102
post #48

Earlier quoted context omitted.

>> "training data is totally a violation of copyright" > This really isn't clear because cognition is treated as a special exception to copyright. Human cognition; not the latest algorithms and their output, which some enthusiastic software engineers eagerly confuse for cognition. It's actually pretty clear.

As I said, human cognition is a special case. The open question is how to handle machines that mimic the process.

> The open question is how to handle machines that mimic the process.

It's not really an open question, except for software engineers who've talked themselves into thinking of humans as computers. A machine is not a human mind, so does not benefit from the legal exceptions and rights granted to the latter.

Re: Judge rejects most ChatGPT copyright claims from book authors

#103
post #36

I think the real issue here is that we’re trying to apply laws made by humans for humans to something entirely new and not well understood: LLMs. This won’t be something we’ll cleanly solve using current regulation.

I agree, but of course national legislation is like impossible now due to a certain obstructionist political party.

You should be more clear about the topics of legislation in question.

In matters involving internet service providers in the US, the Democratic party in Congress generally doesn't bother to propose internet laws either way, but the Republican party tends to strip away the FCC's regulatory authority and actively oppose attempts to bring a fraction back [1][2].

[1] https://arstechnica.com/tech-policy/2023/12/ted-cruz-wants-t...

[2] https://arstechnica.com/tech-policy/2024/01/fcc-to-halt-sign...

Is either party attempting to obstruct regulations on AI training or AI usage? At the moment, I don't think so [3].

[3] https://www.politico.com/news/2023/09/13/schumer-senate-ai-p...

Re: Judge rejects most ChatGPT copyright claims from book authors

#104

Though I think that training data is totally a violation of copyright, OpenAI really needs to win. Copyright has been unreasonably extended to the point where it's untenable and if we needed to track down rights holders and negotiate 'training rights' we'd ensure there would be no open models or competition going forward. The rights holders are going to lose this one and it's probably for the better.

> Though I think that training data is totally a violation of copyright, OpenAI really needs to win.

I'm not sure how to interpret this line. By "violation of copyright" are you referring to the ideal copyright law that's in your mind but doesn't exist yet, or are you referring to the copyright statutes and case law that are currently in effect? And in your opinion is OpenAI violating the latter, the former, or both?

Re: Judge rejects most ChatGPT copyright claims from book authors

#105
post #74

Earlier quoted context omitted.

Large language models are also not databases of text.

Thought question, not entirely related but if you want to go that that route it actually is. If I generate some media in say Photoshop. I then send you a JPEG representation of said media. You then distribute a PNG copy of the image without license. Have you violated copyright law? At what point is there enough parameters to an LLM that it is effectively just a compressed version. How about deduplicated storage? Is a…

The human brain has 100 billion neurons, is it just effectively compressing everything that it's ever seen? I don't actually know but my feeling is no.

If I ask you to draw Mickey Mouse, you can probably produce a very good representation of him. If I asked you to write the script of The Matrix, assuming you've seen it, I suspect you'd get all the plot points down and major quotes even if it has been years since you've seen it. Are you creating a copy? Absolutely! Don't distribute either of those things without a license. But does the fact that you are capable of making a copy of a thing when asked mean that you've violated copyright way back when you watched the Matrix? Is there a copy of the Matrix or Mickey Mouse in your brain?

I will take the strong position that neither our brains nor LLMs contain copies of data in the way that is a violation of copyright. But both are equally capable of generating copyright violating materials.

Re: Judge rejects most ChatGPT copyright claims from book authors

#106

Earlier quoted context omitted.

> You're conflating different matters here by merging together "legal/illegal" and "successfully prosecuted" There isn't any difference between these two things. If something has never and will never be enforced then it is just words on a piece of paper. The only way that words on a piece of paper turn into something that actually matters is through enforcement. So the point stands. It doesn't matter what your interp…

"not prosecuted" doesn't mean "not enforced", it means "not enforced by state ". Companies can successfully enforce copyright infringement and sometimes do so - there's a world of difference "if I don't attract too much attention, it's very unlikely they'll come after me" and "what I'm doing is legal".

> it means "not enforced by state

Ok got it.

Well then you can reinterprete everything that I said previously to instead be "nobody has either been prosecuted or sued successfully by either the state or civility for the downloading part".

I didn't realize that the problem that you had with my post was not with the clear substance of it, and was instead with the dictionary definition and usage of one single word.

But if that one singular word was the issue then I am happy to give this multi sentence clarification even though I think my point was clear and obvious from the beginning.

So my original point stands once I have corrected the slightly incorrect usage of one single word.

But feel free to show an example of a single person ever being successfully sued for just the "downloading" part.

(civilly! Not by the state instead by a company! Important clarification here. I don't want you to misinterpret. Because apparently that one word was a big deal and huge misunderstanding.)

Re: Judge rejects most ChatGPT copyright claims from book authors

#107
post #74

Earlier quoted context omitted.

Large language models are also not databases of text.

Thought question, not entirely related but if you want to go that that route it actually is. If I generate some media in say Photoshop. I then send you a JPEG representation of said media. You then distribute a PNG copy of the image without license. Have you violated copyright law? At what point is there enough parameters to an LLM that it is effectively just a compressed version. How about deduplicated storage? Is a…

> I then send you a JPEG representation of said media. You then distribute a PNG copy of the image without license.

The purpose and function of an image format is to represent a single piece. The representation can vary in accuracy and can be changed to another representation (JPEG to PNG, or one JPEG implementation to another JPEG implementation), but the underlying piece is supposed to be the same in intent and a majority of the time is the same in practice. Open the JPEG image using any program that implements the JPEG specification and you will get the same image as with a different such program. The same would apply to the use of encryption on the image, but not to a cryptographic hash (designed to be one-way). Decrypt the encrypted image and you'll get something that's the same in intent and in technicality as the original image. If you don't use a tool which has the purpose of decrypting things then you almost certainly won't get the same image back. The same only partially but at least partially applies to an AI: if you don't try to use the AI to reproduce an existing work then you still might get a reproduction of an existing work; the probability of such a result varies greatly depending on the prompt.

The purpose of an LLM isn't to reproduce one or more works - or rather, sections of expression - in the dataset. The purpose of an LLM is to produce speech similar to a human's response. The purpose of an image generator model is to produce images that have the characteristics specified in the prompt. In order to produce a copy of something in the training set, the prompt usually needs to reference a specific work, a related person (e.g. an author), a related work, or an attribute that is strongly associated with a particular work/author. Regarding the latter, there was a Hacker News post (that I can't find because I forgot the post title) from a month or two back about an AI image generator that produced images of the robot C-3PO from Star Wars even though the prompt was about "space" and "robot" with no reference to Star Wars. My interpretation is that the AI model had a strong association between space robot and Star Wars because C-3PO is (I speculate) one of the most common space-related robots that people talk about online. Or perhaps, the Star Wars works in the training set made up a majority of the works associated with both "space" and "robot". But I digress.

The likelihood that an AI produces a copy of existing expression depends on the prompt. A user who encounters such a case can avoid liability by not using and not sharing the output, and otherwise the output might not substitute for the original expression for the user's purposes. So in most cases I think liability for the infringinging outputs of an AI model should fall solely on the prompter. The liability that should fall on the developer of the AI model doesn't have to be binary. There can be heavier penalties on the developer for an AI that is more likely to reproduce C-3PO when given a vague prompt such as "space robot", lesser penalties on the developer for a model that only produces C-3PO when the prompt is at least as specific as "space war robot", and even lesser or no penalties for a model that only produces C-3PO from a prompt as specific as "golden space robot". The threshold for vague prompt would vary; for a prompt such as "painting of melting clock" I would excuse a partial reproduction of Salvador Dalí's The Persistence of Memory.

Re: Judge rejects most ChatGPT copyright claims from book authors

#108

Though I think that training data is totally a violation of copyright, OpenAI really needs to win. Copyright has been unreasonably extended to the point where it's untenable and if we needed to track down rights holders and negotiate 'training rights' we'd ensure there would be no open models or competition going forward. The rights holders are going to lose this one and it's probably for the better.

Copyright laws need to be reformed but I don't understand why we need to just completely upend the concept of Copyright just because it's inconvenient to these massive tech companies that want to hoover up all the data they can with as little interference and COST to them as possible. Let's not pretend these are some small startups by two employees working out of a garage who took out reverse mortgages on their homes.

I can see the argument it might reduce open models, but how is asking for permission (and maybe sometimes paying) of the rights holders going to reduce competition? It would be a level playing field. If you're a company that's working on a LLM, you have to ask permission to ingest data that doesn't belong to you. All companies working on an LLM would have to follow these rules. So how is it anti-competitive?

Re: Judge rejects most ChatGPT copyright claims from book authors

#109
post #67

Earlier quoted context omitted.

And we won't have open models.

Try training an "open" model on Nintendo's, Disney's, and Elsevier's IP and see how long it takes them to bury you in lawsuits citing copyright infringement. The only way out of this would be to abolish copyright.

Or indeed carve out an exception for AI training in recognition for their potential to revolutionize our society.

Re: Judge rejects most ChatGPT copyright claims from book authors

#110

Though I think that training data is totally a violation of copyright, OpenAI really needs to win. Copyright has been unreasonably extended to the point where it's untenable and if we needed to track down rights holders and negotiate 'training rights' we'd ensure there would be no open models or competition going forward. The rights holders are going to lose this one and it's probably for the better.

> Though I think that training data is totally a violation of copyright, OpenAI really needs to win. I’d argue that that’s not really how the law is supposed to work.

Isn’t that what “jury nullification” is? Deciding that a law is trash and you’re just going to do whatever instead? I realise this requires a jury which isn’t the case for everything.
Post reply on HN