Live data from Hacker News

We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

stablediffusionlitigation.com

301–310 of 473 posts

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#301
post #2

“Sta­ble Dif­fu­sion con­tains unau­tho­rized copies of mil­lions—and pos­si­bly bil­lions—of copy­righted images.” That’s going to be hard to argue. Where are the copies? “Hav­ing copied the five bil­lion images—with­out the con­sent of the orig­i­nal artists—Sta­ble Dif­fu­sion relies on a math­e­mat­i­cal process called dif­fu­sion to store com­pressed copies of these train­ing images, which in turn are recom­bine…

One idea I had was to try to recreate the original using a prompt. If you succeed, it should be obvious that the original was in the training set?

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#302
post #225

Earlier quoted context omitted.

Pretty sure that’s already decided. Publicly played movies and music are not available to be used. Why would the same not apply to posted images?

If you post a song on your website and I listen to it am I violating your copyright? If my parrot recites your song after hearing my alleged infringement, I record its performance and post it on YouTube is that infringement? Last one, if I use the song from your website to train an song recognition AI is that infringement?

If I host a song I don't have license to on my website I'm violating copyright by distributing it to you when you listen on my site.

If my parrot recites your song after hearing it and I record that and upload to YouTube. I've violated your copyright.

If a big company does the same(runs the song through a non-human process, then sells the output) I believe they're blatantly infringing copyright.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#303

Earlier quoted context omitted.

Great. Now the defence shows an artist that can recreate an image. Cool, now people who look at images get copyright suits filed against them for encoding those images in their heads.

Just because I look at an image does not mean that I can recreate it. storing it in the training data means the AI can recreate it. There's a world of difference that you are just writing off.

If you spent a decade trying to draw it, wouldn't your brain have the right "weights" to execute it pretty exactly going forward?

Except with computers, they don't need to eat or sleep, converse or attend stand-ups.

And once you're able to draw that one picture, you could probably draw similar ones. Your own style may emerge too.

Just thinking. Copywriters, students, and scribes used to copy stuff verbatim, sometimes just to "learn" it.

The product of that study could be published works, a synthesis of ideas from elsewhere, and so on. We would say it belonged to the executor, though.

So the AI learned, and what it has created belongs to it. Maybe.

Or, once we acknowledge AI can "see" images, precedent opens the way to citizenship (humanship?)

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#304

Earlier quoted context omitted.

Legal

Can you clarify? My understanding is that it's very unclear whether there are any legal issues (in most jurisdictions) in scraping for training. Obviously some fairy reputable organisations and individuals are moderately confident that there isn't otherwise they wouldn't have done it.

"It's very unclear" in legal cases is synonymous with "it hasn't been challenged in court yet". You say they're moderately confident because they're fairly reputable, but remember that Madoff was a "reputable business man" for the 20 years he ran a ponzi scheme. They don't have to be confident in the legality to do it, they just had to be confident in the potential profit. With openai being values at $10B by Microsoft, I'd say they've successfully muddied the legal waters long enough to cash out.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#305
post #271

Earlier quoted context omitted.

What court case set the president that you can’t train a neural network on publicly posted movies and audio?

I'd assume the precedent would be about sharing encoder data, which would be covered in bittorrent cases. "Training a neural network" is an implementation detail. These companies accessed millions of copyrighted works, encoded them such that the copyright was unenforcable, then sell the output of that transformation.

Not being able to reproduce the inputs (each image is contributing single bytes to the neural network) is relevant. Torrent files are a means to exactly reproduce their inputs. Diffusion models are trained to not reproduce their inputs, nor do they have the means to.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#306

Earlier quoted context omitted.

If you post a song on your website and I listen to it am I violating your copyright? If my parrot recites your song after hearing my alleged infringement, I record its performance and post it on YouTube is that infringement? Last one, if I use the song from your website to train an song recognition AI is that infringement?

If I host a song I don't have license to on my website I'm violating copyright by distributing it to you when you listen on my site. If my parrot recites your song after hearing it and I record that and upload to YouTube. I've violated your copyright. If a big company does the same(runs the song through a non-human process, then sells the output) I believe they're blatantly infringing copyright.

Big Company is not distributing the input images by distributing the neural network. There is no way to extract even a single input image out of a diffusion model.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#307
post #26
post #20

Earlier quoted context omitted.

I don't think you have to reproduce an entire original work to demonstrate copyright violation. Think about sampling in hip hop for example. A 2 second sample, distorted, re-pitched, etc. can be grounds for a copyright violation.

The difference here is that the images aren't stored, but rather an extremely abstract description of the image was used to very slightly adjust a network of millions of nodes in a tiny direction. No semblance of the original image even remotely exists in the model.

> No semblance of the original image even remotely exists in the model

What does this mean? It doesn't mean you can't recreate the original, because that's been done. It doesn't mean that literally the bits for the image aren't present in the encoded data, because that's true for any compression algorithm.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#308
post #88
post #44

Earlier quoted context omitted.

You could make the same argument that as long as you are using lossy compression you are unable to infringe on copyright.

That's a huge understatement. 5 billion images to a model of 5GB. 1 byte per image. Let's see if one byte per image would constitute a copyright violation in other fields than neural networks.

You took the images, encoded them in a computer process, and the result is able to reproduce some of those images. I fail to see why the size of the training set in bytes and the size of the model in bytes matters. Especially if, as other commenters have noted, much if the training data is repeated(mentions of thousands of mina Lisa's) so a straight division(training size/parameters size) says nothing about the bytes per copyrighted work.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#310
post #95
post #88

Earlier quoted context omitted.

That's a huge understatement. 5 billion images to a model of 5GB. 1 byte per image. Let's see if one byte per image would constitute a copyright violation in other fields than neural networks.

It will be interesting to see how they legally define the moment where compression stops being compression and starts being an original work. If I train on one image I can get it right back out. Even two, maybe even a thousand? Not sure what the line would be where it becomes ok vs not but there will have to be some answer.

There only needs to be an answer if it's determined that some number isn't copyright infringement. The easy answer would be to say that the process is what prevents the works from being transformative(and thus copyrightable) and not the size of the training set.
Post reply on HN