Live data from Hacker News

Judge rejects most ChatGPT copyright claims from book authors

arstechnica.com

111–120 of 126 posts

Re: Judge rejects most ChatGPT copyright claims from book authors

#111

Earlier quoted context omitted.

"not prosecuted" doesn't mean "not enforced", it means "not enforced by state ". Companies can successfully enforce copyright infringement and sometimes do so - there's a world of difference "if I don't attract too much attention, it's very unlikely they'll come after me" and "what I'm doing is legal".

> it means "not enforced by state Ok got it. Well then you can reinterprete everything that I said previously to instead be "nobody has either been prosecuted or sued successfully by either the state or civility for the downloading part". I didn't realize that the problem that you had with my post was not with the clear substance of it, and was instead with the dictionary definition and usage of one single word. But…

Sure, it was quite widespread in early 2000s, with tens of thousands people sued. While the main poster cases like the $675000 award of Sony vs Tenenbaum also included assertions of distribution (AFAIK not proving any specific upload, but only "making available"), there are cases where only dowloads were asserted such as Cassi Hunt (https://www.acslaw.org/?post_type=acsblog&p=3005) and others, and the key part is the many thousands of people who settled for ~$3000 each, which I'd count as "successfully suing" even if the case never went to a court, as that acts as mass enforcement through civil means that has at least some impact on how people behave.

Re: Judge rejects most ChatGPT copyright claims from book authors

#112
post #61

Earlier quoted context omitted.

Fair use is non-transitive: you reviewing a pirated copy of a movie can be fair use even if that copy isn't. If training on copyrighted images is fair use then it doesn't matter how you got those images. Think about it this way: if the opposite were true, then being able to review a movie would be a privilege you have to pay for by buying the movie, rather than just something you can do because the 1st Amendment exis…

It has not been proven that AI training data is fair use.

Pretend my comment starts with "Assuming AI training is ruled to be fair use".

I don't actually think it will be ruled fair use, at least not in all cases.

Re: Judge rejects most ChatGPT copyright claims from book authors

#113
post #9

Earlier quoted context omitted.

Yeah agreed, and also the general direction this is heading in feels a bit strange to me - we can debate the merits of current copyright law - but the indisputable fact is that the training material came from *pirated* copies of these author’s copyrighted works. If we are to believe the much used arguments against piracy that we’ve been fed over the last 30+ years (mainly by deep pocketed media companies) then piracy…

The piracy argument can be fixed by OpenAI buying one copy of each work. The overall question of whether they're allowed to train on copyrighted material without permission seems much larger and more interesting.

Exactly! They just need to pay for it!

The same way that they should pay for using news articles and people’s images/comics.

Re: Judge rejects most ChatGPT copyright claims from book authors

#114
post #74

Earlier quoted context omitted.

Large language models are also not databases of text.

Thought question, not entirely related but if you want to go that that route it actually is. If I generate some media in say Photoshop. I then send you a JPEG representation of said media. You then distribute a PNG copy of the image without license. Have you violated copyright law? At what point is there enough parameters to an LLM that it is effectively just a compressed version. How about deduplicated storage? Is a…

Of course it matters. If I put the right data into a paintbrush and canvas I'll reproduce copyrighted works too. Nobody is confusing the model for a Picasso anymore than they confuse a paintbrush for a painting. The law may rule differently for one reason or another, but these are obviously different categories of things.

The reason a jpeg is copyright infringement has more to do with its express purpose being to allow the user to view that copyrighted work. If it were bundled in a program that just allowed you to view the color histogram of famous works (and the author had the right to view those photos and didn't think it important to save bandwidth by precomputing those histograms) it likely wouldn't be infringement. If it were found out that people were downloading that program just to rip the bundled images out then the author might get in hot water anyway. Your distributed file example is similarly probably infringing.

The model has other capabilities, and I think it would be hard to argue that its purpose is copyright infringement (which is separate from what you seem to be doing, which is arguing that the model itself is infringement -- both a little easier and harder to argue because it pushes more on philosophical distinctions than statements of fact about how people are using a thing).

Separately, there are new classes of concerns these models introduce. We don't have to abuse copyright law to take the time to consider those effects. E.g., should voice cloning be allowed and to what degree? It's already illegal in a lot of contexts (fraud, ...), but we don't currently have many rights when it comes to our innate physical characteristics. To the extent those rights exist, you often have to waive them for basic services (e.g., a nontrivial fraction of leases and jobs stipulate that you give a permanent, , license for them to use your image for nearly any purpose, including falsely characterizing your approval of the property in advertisements and marketing materials -- unless covered under libel/slander and a couple other carveouts they're probably not punishable). Can studios just refuse to hire voice actors for more than one session? Is that good for society? Can I clone passers-by on the street to play in my commercial? These are new enough capabilities (at least at their current scale) that they're not very well legislated, and I wouldn't be surprised if we saw an expansion of something like "moral rights" to cover them.

Re: Judge rejects most ChatGPT copyright claims from book authors

#115
post #102

Earlier quoted context omitted.

As I said, human cognition is a special case. The open question is how to handle machines that mimic the process.

> The open question is how to handle machines that mimic the process. It's not really an open question, except for software engineers who've talked themselves into thinking of humans as computers. A machine is not a human mind, so does not benefit from the legal exceptions and rights granted to the latter.

This has nothing to do with how things operate, or whether an LLM is like a mind. It's a legal question regarding large scale compilation of data.

"A machine is not a human mind, so does not benefit from the legal exceptions and rights granted to the latter."

Five years ago this was true. It likely will not be true, eventually. The only question is where we are in this process.

Re: Judge rejects most ChatGPT copyright claims from book authors

#116
post #36

Earlier quoted context omitted.

I agree, but of course national legislation is like impossible now due to a certain obstructionist political party.

[assuming you mean America] as long as there're only two parties they're both equally obstructionist.

I do and that's wrong.

Re: Judge rejects most ChatGPT copyright claims from book authors

#117
post #74

Earlier quoted context omitted.

Thought question, not entirely related but if you want to go that that route it actually is. If I generate some media in say Photoshop. I then send you a JPEG representation of said media. You then distribute a PNG copy of the image without license. Have you violated copyright law? At what point is there enough parameters to an LLM that it is effectively just a compressed version. How about deduplicated storage? Is a…

The human brain has 100 billion neurons, is it just effectively compressing everything that it's ever seen? I don't actually know but my feeling is no. If I ask you to draw Mickey Mouse, you can probably produce a very good representation of him. If I asked you to write the script of The Matrix, assuming you've seen it, I suspect you'd get all the plot points down and major quotes even if it has been years since you'…

However, one can memorize something like a book, particularly if one uses known "memory palace" techniques. Some individuals are particularly good at this.

Re: Judge rejects most ChatGPT copyright claims from book authors

#118

Earlier quoted context omitted.

> it means "not enforced by state Ok got it. Well then you can reinterprete everything that I said previously to instead be "nobody has either been prosecuted or sued successfully by either the state or civility for the downloading part". I didn't realize that the problem that you had with my post was not with the clear substance of it, and was instead with the dictionary definition and usage of one single word. But…

Sure, it was quite widespread in early 2000s, with tens of thousands people sued. While the main poster cases like the $675000 award of Sony vs Tenenbaum also included assertions of distribution (AFAIK not proving any specific upload, but only "making available"), there are cases where only dowloads were asserted such as Cassi Hunt ( https://www.acslaw.org/?post_type=acsblog&p=3005 ) and others, and the key part is t…

> people who settled for ~$3000 each

So then not a single instance of an actual judge ruling that the downloading part was civilly illegal or requires payment civilly?

If no, then my point stands.

Anyone can make a threat and get money out of someone.

I could threaten you right now with horrible legal consequences for the completely bogus "civil crime" of being wrong on hacker news!

But you giving in to me, and giving me money, isn't a ruling or precedent that action X (in this case "being wrong on hacker news") is illegal.

Re: Judge rejects most ChatGPT copyright claims from book authors

#119
post #9

Earlier quoted context omitted.

The piracy argument can be fixed by OpenAI buying one copy of each work. The overall question of whether they're allowed to train on copyrighted material without permission seems much larger and more interesting.

Exactly! They just need to pay for it! The same way that they should pay for using news articles and people’s images/comics.

There are two versions of "just pay for it" though:

1. Pay retail price for one copy then train on it.

2. Pay a recurring license fee that probably amounts to the majority of your revenue (Spotify-style).

AI companies can't really afford #2.

Re: Judge rejects most ChatGPT copyright claims from book authors

#120

Earlier quoted context omitted.

Sure, it was quite widespread in early 2000s, with tens of thousands people sued. While the main poster cases like the $675000 award of Sony vs Tenenbaum also included assertions of distribution (AFAIK not proving any specific upload, but only "making available"), there are cases where only dowloads were asserted such as Cassi Hunt ( https://www.acslaw.org/?post_type=acsblog&p=3005 ) and others, and the key part is t…

> people who settled for ~$3000 each So then not a single instance of an actual judge ruling that the downloading part was civilly illegal or requires payment civilly? If no, then my point stands. Anyone can make a threat and get money out of someone. I could threaten you right now with horrible legal consequences for the completely bogus "civil crime" of being wrong on hacker news! But you giving in to me, and givin…

All the trial cases which went through the courts e.g. Sony vs Tenenbaum also were judged that the downloading part was also infringing activity, it's just that those defendants did both downloading and [offering for] distribution.
Post reply on HN