Live data from Hacker News

We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

stablediffusionlitigation.com

361–370 of 473 posts

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#361

It seems to me the communal voice of HN varies widely on copyright issues depending on who is getting sued and who is getting potentially hurt by violations. People who generally make less money than programmers - writers, artists, musicians - should stop their whining and their unfair uses of copyright to control their creative output. Programmers who are getting shafted by big corporations using their code to build…

There is no inconsistency. The programmers suing Github put their code online purposely for other people to use in their own open source software for free. They aren't suing for their protection; quite the opposite. The programmers are suing out of a hatred of copyright itself, not an interest in getting themselves better copyright protection for their work, nor for money.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#362

So what is the end goal of this? For copyright to transfer every step? That precedent happens, then what? Licensing schemes get set up and any piece of media that is put into these systems will result in the artist getting some kind of payment in return. Cool, that sound great. Except... who's paying? The conglomerates who already have a bunch of IP they can feed into those systems, who can afford to purchase or thro…

You're arguing that artists have a shitty home, therefore it's not worth protecting as those AI companies are trying to take even that from them. And you're somehow trying to sound like you're pro artists in all this. Please, listen to yourself.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#363

Earlier quoted context omitted.

This is very much a 'color of your bits' topic, but I'm not sure why the internal representation matters. It's pretty trivial to recreate famous works like the Mona Lisa or Starry Night or Monet's Water Lily Pond. Obviously some representation of the originals exist inside the model+prompt. Why wouldn't that apply to other images in the training sets?

>It's pretty trivial to recreate famous works like the Mona Lisa or Starry Night or Monet's Water Lily Pond. A recreation of a piece of art does not mean a copy, I've personally seen hundreds of recreations of Edvard Munch's 'The Scream', all of them perfectly legal. Even in a massively overtrained model, it is practically impossible to create a 1:1 copy of a piece of art the model was trained upon. And of course tha…

A work doesn't have to be identical to be considered a derivative work, which is why we also don't consider every JPEG a newly copyrighted image distinct from the source material.

As an example of a plausible scenario where copyright might actually be violated, consider this: an NGO wants images on their website. They type in something like 'afghan girl' or 'struggling child' and unknowingly use the recreations of the famous photographs they get.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#364

Earlier quoted context omitted.

Just because it generates you an image like Biden still does not make it a derivative either. You can draw Biden yourself if you're talented and it's not considered a derivative of anything.

The difference is that computers create perfect copies of images by default, people don't. If a person creates a perfect copy of something it shows they have put thousands of hours of practice into training their skills and maybe dozens or even hundreds of hours into the replica. When a computer generates a replica of something it's what it was designed to do. AI art is trying to replicate the human process, but it w…

>The difference is that computers create perfect copies of images by default

are we looking at the output of the same program? because all of the output images i look at have eyes looking in different direction and things of horror in place of hands or ears, and they feature glasses meting into people faces, and that's the good ones, the bad one have multiple arms contorting out of odd places while bent at unnatural angles.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#365

Earlier quoted context omitted.

So as a code author I am pretty upset about Copilot specifically, and it seems like SD is similar (hadn't heard before about DeviantArt doing the same as what GitHub did). But I agree with this take: the tech is here, it's going to be used, and it's not going to be shut down by a lawsuit. Nor should it, frankly. What I object to is not the AI itself, or even that my code has been used to train it. It's the copyright…

> Have they actually trained Copilot on their own source? If not, why not? People have posted illegal Windows source code leaks to GitHub. Microsoft doesn’t seem to care that much because these repos stay up for months or even years at a time without Microsoft DMCAing them-if you go looking you’ll find some right now. I think it is entirely possible, even likely, that some of those repos were included in Copilot’s tr…

The question is not whether there's some of their code that they don't mind being incorporated, but whether there's any at all that they wouldn't allow to be. And more importantly, not used for their own bot, but for someone else's.

If licenses don't apply to training, then they don't apply for anyone, anywhere. If they do apply, then Copilot is violating my license.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#366
post #42

Earlier quoted context omitted.

> 90%ish of a single input image Oh, one image is enough to apply copyright as if it were a patent, to ban a process that makes original works most of the time? The article authors say it works as a "collage tool" trying to minimise the composition and layout of the image as unimportant elements. At the same time forgetting that SD is changing textures as well, so it's a collage minus textures and composition? Is the…

But they are not original works, they are wholly derived works of the training data set. Take that data set away and the algorithm is unable to produce a single original pixel. The fact that the derivation involves millions of works as opposed to a single one is immaterial for the copyright issue.

how is that any different from new human artist that study other artists work to learn a style or technique. In fact it used to be that the preferred way for painters to learn was to repeatedly copy paintings of masters.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#367

Earlier quoted context omitted.

This surely can't be the case, right? If it was, then what's stopping me from taking any possible byte sequence and applying my copyright to it? I could always show that there exists some function f that produces said byte sequence when applied to my copyrighted material. Can I sue Microsoft because the entire Windows 11 codebase is just one "rote mathematical transformation" away from the essay I wrote in elementary…

The law doesn't care about technical tricks. It cares about how you got the bytes and what humans think of them. Sure, the windows 11 codebase is in pi somewhere if you go far enough. Sure, pi is a non-copyrightable fact of nature. That doesn't mean the windows codebase is _actually_ in pi legally, just that it technically is. The law does not care about weird gotchas like you describe. I recommended reading this to…

>But, it's still copyrighted by the original artist as long as they can show "This started as my image, and a machine made a rote mathematical transformation to it"

I think the post you’re replying to saw was confused about the quote above. The person who’s claiming copyright by showing the claimed file started as their own image has to show that it started from their own image, and not just that the file could have derived from the image. Copyright cares about both the works and the provenance of works.

Stable Diffusion couldn’t be flagged under this pretense if a person used a prompt that was their own nor could they even be sued if they ran an image through it as long as there is no plausibility that it was made by a copyright work. The only thing I imagine a case working on is the actual training process of the algorithm rather than the algorithm itself for that exact reason.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#368

Earlier quoted context omitted.

Can you clarify? My understanding is that it's very unclear whether there are any legal issues (in most jurisdictions) in scraping for training. Obviously some fairy reputable organisations and individuals are moderately confident that there isn't otherwise they wouldn't have done it.

"It's very unclear" in legal cases is synonymous with "it hasn't been challenged in court yet". You say they're moderately confident because they're fairly reputable, but remember that Madoff was a "reputable business man" for the 20 years he ran a ponzi scheme. They don't have to be confident in the legality to do it, they just had to be confident in the potential profit. With openai being values at $10B by Microsof…

That's one company. There's dozens if not hundreds of companies, research groups and individuals working under the same assumption.

Maybe it's a mass delusion but that feels like a stretch.

Also your wording makes this sound entirely like a sinister conspiracy or cash grab. Many people think this is simply a worthy pursuit and the right direction to be looking at the moment.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#369

Earlier quoted context omitted.

The difference is that computers create perfect copies of images by default, people don't. If a person creates a perfect copy of something it shows they have put thousands of hours of practice into training their skills and maybe dozens or even hundreds of hours into the replica. When a computer generates a replica of something it's what it was designed to do. AI art is trying to replicate the human process, but it w…

>The difference is that computers create perfect copies of images by default are we looking at the output of the same program? because all of the output images i look at have eyes looking in different direction and things of horror in place of hands or ears, and they feature glasses meting into people faces, and that's the good ones, the bad one have multiple arms contorting out of odd places while bent at unnatural…

Storing and retrieving photos, files, music, exactly identical to how they were before, is what computers do.

Save a photo on your computer, open it in a browser or photo viewer, you will get that photo. That is the default behavior of computers. That is not in dispute, is it?

All of this machine learning stuff is trying to get them to not do that. To actually create something new that no one actually stored on them.

Hope that clears up the misunderstanding.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#370
post #26

Earlier quoted context omitted.

The difference here is that the images aren't stored, but rather an extremely abstract description of the image was used to very slightly adjust a network of millions of nodes in a tiny direction. No semblance of the original image even remotely exists in the model.

This is very much a 'color of your bits' topic, but I'm not sure why the internal representation matters. It's pretty trivial to recreate famous works like the Mona Lisa or Starry Night or Monet's Water Lily Pond. Obviously some representation of the originals exist inside the model+prompt. Why wouldn't that apply to other images in the training sets?

It’s not quite a one to one. Copyright law isn’t as arbitrary as it would seem in my experience. Also there’s the conflation of two things here: whether the model is within copyright violation and whether the works generated by it are

The “color of your bits” only applies to the process of creating a work. Stable Diffusion’s training of the algorithm could be seen as violating copyright but that doesn’t spread to the works generated by it.

In the same vein, one can claim copyright on an image generated by stable diffusion even if the creation of the algorithm is safe from copyright violation.

“some representation of the originals exist inside the model+prompt” is also not sufficient for the model to be in violation of copyright of any one art piece. Some latent representation of the concept of an art piece or style isn’t enough.

It’s also important to note the distinction that there is no training data stored in its original form as part of the model during training, it’s simply used to tweak a function with the purpose of translating text to images. Some could say that’s like using the color from a picture of a car on the internet. Some might say it’s worse but it’s all subjective unless the opposition can draw new ties of the actual technical process to things already precedent.

Post reply on HN