Live data from Hacker News

We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

stablediffusionlitigation.com

271–280 of 473 posts

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#271
post #225
post #113

Earlier quoted context omitted.

“Did they have a right to use publicly posted images” is up for the courts to decide

Pretty sure that’s already decided. Publicly played movies and music are not available to be used. Why would the same not apply to posted images?

What court case set the president that you can’t train a neural network on publicly posted movies and audio?

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#272
post #201
post #46

Earlier quoted context omitted.

Every professional artist trained their own brains on some number copyrighted images, without the consent of the original creator.

I’m aware of that argument, but it's not a silver bullet. Scale matters. There’s, for instance, a difference between looking at your neighbors house vs recording hi-res video. Human attention is an incredibly scarce resource, and arguably even sacred.

A neural network is anything but a high-res recording of the content it’s trained on. It’s got a particularly low resolution view (512x512 iirc) of its inputs.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#273
post #42

Earlier quoted context omitted.

> 90%ish of a single input image Oh, one image is enough to apply copyright as if it were a patent, to ban a process that makes original works most of the time? The article authors say it works as a "collage tool" trying to minimise the composition and layout of the image as unimportant elements. At the same time forgetting that SD is changing textures as well, so it's a collage minus textures and composition? Is the…

But they are not original works, they are wholly derived works of the training data set. Take that data set away and the algorithm is unable to produce a single original pixel. The fact that the derivation involves millions of works as opposed to a single one is immaterial for the copyright issue.

If I take a million copywritten images from magazines, cut them with scissors, and make a single collage, I would expect the resulting image to be fair use. Fair use is an affirmative defense, like self defense, where you justify your infringement.

People are treating this like its a binary technical decision. Either it is or isn't a violation. Reality is that things are spectrums and judges judge. SD will likely be treated like a remix that sampled copywritten work, but just a tiny bit of each work, and sufficiently transformed it to create a new work.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#274

Earlier quoted context omitted.

People in art school also practice by studying existing art and images.

I think what’s clear is that this is an unprecedented type of use. I’m really interested in seeing how the courts rule on this one as it has wide implications for the AI era.

Because this use is unprecedented, as you say, it's clear that the law wasn't written with this use case in mind. The more interesting question in my mind is what we think the new law should be, rather than what the courts happen to make of the existing law. I.e., I think the answer should come from politicians not from judges.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#275
post #149

Earlier quoted context omitted.

Just because I look at an image does not mean that I can recreate it. storing it in the training data means the AI can recreate it. There's a world of difference that you are just writing off.

No, it means there is a 512 bit number you can combine with the training data to reproduce a reasonable though not exact likeness (attempts to use SD and others as compression algorithms show they're pretty bad at it, because while they can get "similar" they'll outright confabulate details in a plausible looking way - i.e. redrawing the streets of San Francisco in images of the golden gate bridge). Which of course t…

> It's equivalent to trying to sue a compression codec because a specific archive contains a copyrighted image.

That's plainly untrue, as Stable Diffusion is not just the algorithm, but the trained model—trained on millions of copyrighted images.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#276

Earlier quoted context omitted.

If I were to take the first word from a thousand books and use it to write my own would I be guilty of copyright violations?

Words have a special carve out in copyright law / precedent. So much so that a whole other category of Intellectual Property exists called Trademarks to protect special words. But back to your point “if you were to take the first sentence from a thousand books and use it in your own book”, then yes based on my understanding (I am not a lawyer) of copyright you would be in violation of IP laws.

I doubt it would be a violation.

Specifically fair use #3 "the amount and substantiality of the portion used in relation to the copyrighted work as a whole."

A sentence being a copyright violation would make every book review in the world illegal.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#277

Why do we keep posting the same arguments over and over with these stories. It's like how humans learn. It contains chunks of copywriter material. What about copilot. Hackers don't respect artists. Yadda yadda. It's boring. I don't know the answer, but after reading the same things over and over I don't know if I trust myself to even have a valid opinion about it.

There’s a pretty easy answer here actually: if you want to include data in a set of training data for an AI system, you need to have (formal, statutory) permission to use it.

No, you don’t

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#280
post #26

Earlier quoted context omitted.

The difference here is that the images aren't stored, but rather an extremely abstract description of the image was used to very slightly adjust a network of millions of nodes in a tiny direction. No semblance of the original image even remotely exists in the model.

This is very much a 'color of your bits' topic, but I'm not sure why the internal representation matters. It's pretty trivial to recreate famous works like the Mona Lisa or Starry Night or Monet's Water Lily Pond. Obviously some representation of the originals exist inside the model+prompt. Why wouldn't that apply to other images in the training sets?

>It's pretty trivial to recreate famous works like the Mona Lisa or Starry Night or Monet's Water Lily Pond.

A recreation of a piece of art does not mean a copy, I've personally seen hundreds of recreations of Edvard Munch's 'The Scream', all of them perfectly legal.

Even in a massively overtrained model, it is practically impossible to create a 1:1 copy of a piece of art the model was trained upon.

And of course that would be a pointless exercise to begin with, why would anyone want to generate 1:1 copies (or anything near that) of existing images ?

The whole 'magic' of Stable Diffusion is that you can create new works of art in the combined styles of art, photography etc that it has been trained on.

Post reply on HN