Live data from Hacker News

We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

stablediffusionlitigation.com

331–340 of 473 posts

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#331
post #26

Earlier quoted context omitted.

The difference here is that the images aren't stored, but rather an extremely abstract description of the image was used to very slightly adjust a network of millions of nodes in a tiny direction. No semblance of the original image even remotely exists in the model.

> No semblance of the original image even remotely exists in the model What does this mean? It doesn't mean you can't recreate the original, because that's been done. It doesn't mean that literally the bits for the image aren't present in the encoded data, because that's true for any compression algorithm.

Do you have any examples of recreating an image with these models? Something other than Mona lisa or other famous artworks because they have caused over fitting.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#332
post #301
post #2

“Sta­ble Dif­fu­sion con­tains unau­tho­rized copies of mil­lions—and pos­si­bly bil­lions—of copy­righted images.” That’s going to be hard to argue. Where are the copies? “Hav­ing copied the five bil­lion images—with­out the con­sent of the orig­i­nal artists—Sta­ble Dif­fu­sion relies on a math­e­mat­i­cal process called dif­fu­sion to store com­pressed copies of these train­ing images, which in turn are recom­bine…

One idea I had was to try to recreate the original using a prompt. If you succeed, it should be obvious that the original was in the training set?

The LAION-5B dataset is public, so you can check directly whether a picture is in there or not. StabilityAI only takes a very limited amount of information from each individual picture, so for Stable Diffusion to closely reproduce a picture it would need to appear quite frequently in the dataset. There are examples of this, such as old famous paintings, "bloodborne box art" and probably many others, though I haven't looked deeply into it.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#333
post #149

Earlier quoted context omitted.

No, it means there is a 512 bit number you can combine with the training data to reproduce a reasonable though not exact likeness (attempts to use SD and others as compression algorithms show they're pretty bad at it, because while they can get "similar" they'll outright confabulate details in a plausible looking way - i.e. redrawing the streets of San Francisco in images of the golden gate bridge). Which of course t…

> It's equivalent to trying to sue a compression codec because a specific archive contains a copyrighted image. That's plainly untrue, as Stable Diffusion is not just the algorithm, but the trained model—trained on millions of copyrighted images.

But in fairness, even a human could know how to violate copyright but cannot be sued until they do violate it.

SD might know how to violate copyright but is that enough to sue it? Or can you only sue violations it helps create?

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#335
post #301
post #2

“Sta­ble Dif­fu­sion con­tains unau­tho­rized copies of mil­lions—and pos­si­bly bil­lions—of copy­righted images.” That’s going to be hard to argue. Where are the copies? “Hav­ing copied the five bil­lion images—with­out the con­sent of the orig­i­nal artists—Sta­ble Dif­fu­sion relies on a math­e­mat­i­cal process called dif­fu­sion to store com­pressed copies of these train­ing images, which in turn are recom­bine…

One idea I had was to try to recreate the original using a prompt. If you succeed, it should be obvious that the original was in the training set?

No, the "original" is in the (likely detailed) prompt you give it.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#336
post #42

Earlier quoted context omitted.

> 90%ish of a single input image Oh, one image is enough to apply copyright as if it were a patent, to ban a process that makes original works most of the time? The article authors say it works as a "collage tool" trying to minimise the composition and layout of the image as unimportant elements. At the same time forgetting that SD is changing textures as well, so it's a collage minus textures and composition? Is the…

But they are not original works, they are wholly derived works of the training data set. Take that data set away and the algorithm is unable to produce a single original pixel. The fact that the derivation involves millions of works as opposed to a single one is immaterial for the copyright issue.

If I make software that randomly draws pixels on the screen then we can say for a fact that no copyrighted images were used.

If that software happens to output an image that is in violation of copyright then it is not the fault of the model. Also, if you ran this software in your home and did nothing with the image, then there's no violation of copyright either. It only becomes an issue when you choose to publish the image.

The key part of copyright is when someone publishes an image as their own. That they copy an image doesn't matter at all. It's what they DO with the image that matters!

The courts will most likely make a similar distinction between the model, the outputs of the model, and when an individual publishes the outputs of the model. This would be that the copyright violation occurs when an individual publishes an image.

Now, if tools like Stable Diffusion are constantly putting users at risk of unknowingly violating copyrights then this tool becomes less appealing. In this case it would make commercial sense to help users know when they are in violation of copyright. It would also make sense to update our copyright catalogues to facilitate these kinds of fingerprints.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#337

Earlier quoted context omitted.

Just because it generates you an image like Biden still does not make it a derivative either. You can draw Biden yourself if you're talented and it's not considered a derivative of anything.

The difference is that computers create perfect copies of images by default, people don't. If a person creates a perfect copy of something it shows they have put thousands of hours of practice into training their skills and maybe dozens or even hundreds of hours into the replica. When a computer generates a replica of something it's what it was designed to do. AI art is trying to replicate the human process, but it w…

You’re acting like the “computer” has a will of it’s own. Generating a perfect copy of an image would be a completely separate task from training a model for image generation.

There are no models I know of with the ability to generate an exact copy of an image from its training set unless it was solely trained on that image to the point it could. In that case I could argue the model’s purpose was to copy that image rather than learn concepts from a broad variety of images to the point it would be almost impossible to generate an exact copy.

I think a lot of the arguments revolving around AI image generators could benefit from the constituent parties reading up on how transformers work. It would at least make the criticisms more pointed and relevant, unlike the criticisms drawn in the linked article.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#338
post #295

Earlier quoted context omitted.

> Oh, one image is enough to apply copyright as if it were a patent, to ban a process that makes original works most of the time? The law can do whatever its writers want. The law is mutable, so the answer to your question is “maybe”. Maybe SD will get outlawed for copyright reasons on a single image. The law and the courts have done sillier things.

All the handwringing about generative AI brings to mind the aphorism about genies returning to bottles. There can be lawsuits and laws--and there may even be cases where an output by chance or by tickling the input sufficiently looks very close to something in the training set. But anyone who thinks this technology will be banned in some manner is... mistaken.

So as a code author I am pretty upset about Copilot specifically, and it seems like SD is similar (hadn't heard before about DeviantArt doing the same as what GitHub did). But I agree with this take: the tech is here, it's going to be used, and it's not going to be shut down by a lawsuit. Nor should it, frankly.

What I object to is not the AI itself, or even that my code has been used to train it. It's the copyright for me but not for thee way that it's been deployed. Does GitHub/Microsoft's assertion that training sidesteps licensing apply to GitHub/Microsoft's own code? Do they want to allow (a hypothetical) FSFPilot to be trained on their proprietary source? Have they actually trained Copilot on their own source? If not, why not?

I published my source subject to a license, and the force of that license is provided by my copyright. I'm happy to find other ways of doing things, but it has to be equitable. I'm not simply ceding my authorship to the latest commercial content grab.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#339
They've got a copy of a figure from the original diffusion paper, showing a diffusion process on a spiral dataset. They seem to completely misunderstand it. The figure does not show image diffusion, rather it shows a diffusion process in which each data item is a 2D point. The figure is showing diffusion on an entire dataset and demonstrating that it can approximately reconstruct the spiral-shaped distribution.

I'm surprised they couldn't find someone with even a rudimentary understanding of diffusion models to review this.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#340
post #12

There are only 2 possible outcomes, right? 1. They win, and big companies are training AI on their portfolio and new artwork which they paid for resulting in the same problem, that artists will even be paid less than today. Also resulting in incredible difficult to judge laws. Has an artist created it himself, or has he used copyrighted works to train his custom model? 2. They lose

Even if they win they lose, because nothing will change. Who is really going to stop using Stable Diffusion at this point?
Post reply on HN