Live data from Hacker News

We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

stablediffusionlitigation.com

291–300 of 473 posts

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#291

Earlier quoted context omitted.

> That’s going to be hard to argue. Where are the copies? In fairness, Diffusion is arguably a very complex entropy coding similar to Arithmetic/Huffman coding. Given that copyright is protectable even on compressed/encrypted files, it seems fair that the “container of compressed bytes” (in this case the Diffusion model) does “contain” the original images no differently than a compressed folder of images contains the…

lol thinking about this more: I understand people’s livelihoods are potentially at stake, but what a shame it would be if we find AGI, even consciousness but have to shut it down because of a copyright dispute.

The real tragedy is being marketed to so heavily that we construe enforcing copyright on llm/diffusion companies with shutting down an AGI. I blame companies like openai purposefully marketing themselves poorly since nobody is going to enforce false advertising laws on something they don't understand.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#292

Earlier quoted context omitted.

> That’s going to be hard to argue. Where are the copies? In fairness, Diffusion is arguably a very complex entropy coding similar to Arithmetic/Huffman coding. Given that copyright is protectable even on compressed/encrypted files, it seems fair that the “container of compressed bytes” (in this case the Diffusion model) does “contain” the original images no differently than a compressed folder of images contains the…

In that vein, surely MD5 hashes should also be copyrighted, as they are derived from a work.

Not really, since one of the major characteristics is being able to recover the copyrighted work from the encoded version.

Since md5 hashes don't share this property, they're not "in that vein".

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#293

Earlier quoted context omitted.

That’s my point, Diffusion[1] does seem to be “just like” gzip or base64. And it would be illegal for me to sell or distribute zipped copies of images without the copyright holder’s consent. Similarly there might be an argument for why Diffusion[1] specifically can’t be built with copyrighted images. [1] which is just one part of something like Stable Diffusion

A lossy compressor isn't just like a lossless compressor. Especially not one that has ~2 bytes for each input image.

So it's fine to distribute copyrighted works, as long as they're jpeg(lossy) encoded? I don't think the law would agree with you.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#294
post #42

Earlier quoted context omitted.

> That’s going to be hard to argue. Where are the copies? In fairness, Diffusion is arguably a very complex entropy coding similar to Arithmetic/Huffman coding. Given that copyright is protectable even on compressed/encrypted files, it seems fair that the “container of compressed bytes” (in this case the Diffusion model) does “contain” the original images no differently than a compressed folder of images contains the…

> 90%ish of a single input image Oh, one image is enough to apply copyright as if it were a patent, to ban a process that makes original works most of the time? The article authors say it works as a "collage tool" trying to minimise the composition and layout of the image as unimportant elements. At the same time forgetting that SD is changing textures as well, so it's a collage minus textures and composition? Is the…

Oh, one image is enough to apply copyright as if it were a patent, to ban a process that makes original works most of the time?

The software itself is not at issue here. If they had trained the network on public domain images then there’d be no lawsuit. The legal question to settle is whether it’s allowable to train (and use) a model on copyrighted images without permission from the artists.

They may actually be successful at arguing that the outputs are either copies or derived works which would require paying the original artist for licenses.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#295
post #42

Earlier quoted context omitted.

> 90%ish of a single input image Oh, one image is enough to apply copyright as if it were a patent, to ban a process that makes original works most of the time? The article authors say it works as a "collage tool" trying to minimise the composition and layout of the image as unimportant elements. At the same time forgetting that SD is changing textures as well, so it's a collage minus textures and composition? Is the…

> Oh, one image is enough to apply copyright as if it were a patent, to ban a process that makes original works most of the time? The law can do whatever its writers want. The law is mutable, so the answer to your question is “maybe”. Maybe SD will get outlawed for copyright reasons on a single image. The law and the courts have done sillier things.

All the handwringing about generative AI brings to mind the aphorism about genies returning to bottles. There can be lawsuits and laws--and there may even be cases where an output by chance or by tickling the input sufficiently looks very close to something in the training set. But anyone who thinks this technology will be banned in some manner is... mistaken.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#296
post #273

Earlier quoted context omitted.

But they are not original works, they are wholly derived works of the training data set. Take that data set away and the algorithm is unable to produce a single original pixel. The fact that the derivation involves millions of works as opposed to a single one is immaterial for the copyright issue.

If I take a million copywritten images from magazines, cut them with scissors, and make a single collage, I would expect the resulting image to be fair use. Fair use is an affirmative defense, like self defense, where you justify your infringement. People are treating this like its a binary technical decision. Either it is or isn't a violation. Reality is that things are spectrums and judges judge. SD will likely be…

If I take a million copywritten images from magazines, cut them with scissors, and make a single collage, I would expect the resulting image to be fair use.

That’s not how it works. Your collage would be fine if it was the only one since you used magazines you bought. Where you’d get into trouble is if you started printing copies of your collage and distributing them. In that case you’d be producing derived works and be on the hook for paying for licenses from the original authors.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#297
post #146

Earlier quoted context omitted.

Go to stablediffusionweb.com and enter "a person like biden" into the box. You will see a picture exactly like President Biden. That picture will have been derived from the trained images of Joe Biden. That cannot be in dispute.

Just because it generates you an image like Biden still does not make it a derivative either. You can draw Biden yourself if you're talented and it's not considered a derivative of anything.

The difference is that computers create perfect copies of images by default, people don't.

If a person creates a perfect copy of something it shows they have put thousands of hours of practice into training their skills and maybe dozens or even hundreds of hours into the replica.

When a computer generates a replica of something it's what it was designed to do. AI art is trying to replicate the human process, but it will always have the stink of "the computer could do this perfectly but we are telling it not to right now"

Take Chess as an example. We have Chess engines that can beat even the best human Chess players very consistently.

But we also have Chess engines designed to play against beginners, or at all levels of Chess play really.

We still have Human-only tournaments. Why? Why not allow a Chess Engine set to perform like a Grandmaster to compete in tournaments?

Because there would always be the suspicion that if it wins, it's because it cheated to play at above it's level when it needed to. Because that's always an option for a computer, to behave like a computer does.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#298
post #271
post #225

Earlier quoted context omitted.

Pretty sure that’s already decided. Publicly played movies and music are not available to be used. Why would the same not apply to posted images?

What court case set the president that you can’t train a neural network on publicly posted movies and audio?

I'd assume the precedent would be about sharing encoder data, which would be covered in bittorrent cases.

"Training a neural network" is an implementation detail. These companies accessed millions of copyrighted works, encoded them such that the copyright was unenforcable, then sell the output of that transformation.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#299
post #262

Earlier quoted context omitted.

I think this is overstretching it. That would be a checksum that can be parsed by humans and contains artistic value that serves as the basis for claims to copyright. An actual checksum no longer has artistic value in itself and cant reproduce the original work. Which is why this is framed as compression, it implies that fundamentally SD makes copies instead of (re)creating art. Leaving out the issue of recreating fo…

Overfitting seems like a fuzzy area here. I could train a model on one image that could consistently produce an output no human could tell apart from the original. And of course, shades of gray from there. Regarding your edit, what are the chances of a "hash collision" where the hash is two MP4 files for two different movies? Seems wildly astronomical.. impossible even? That's why this hash method is so special, plus…

Once you are down to one picture, collisions become feasible given the right environment and resolution of the image.

Pretty sure this is nitpicking about an overused analogy though.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#300

Earlier quoted context omitted.

Just because I look at an image does not mean that I can recreate it. storing it in the training data means the AI can recreate it. There's a world of difference that you are just writing off.

> storing it in the training data means the AI can recreate it. No it doesn't, it means that abstract facts related to this image might be stored.

This just sounds like really fancy, really lossy compression to me.

Compression that returns something different from the original most of the time, but still could return the original.

Post reply on HN