Live data from Hacker News

We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

stablediffusionlitigation.com

441–450 of 473 posts

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#441

Earlier quoted context omitted.

Models for Stable Diffusion are about 2-8GB in size. 5 billion images means that every image gets about 1 byte. It seems to me that they're claiming here that Stability has somehow manage to store copies of these images in about 1 byte of space each. That's an incredible compression ratio!

It is a form compression that loses some much of the uniqueness which gives it the high ratio. If the concept is a little hard to grasp consider an AI model like a finite state machine, but it stores affinity and weights of the data's relationship to each other too. In GPT this is words and phrases, e.g. "Frodo Baggins" high affinity, "Frodo Superman" will be negligible. Now consider all words that may link to those…

I think so too. I also think that this is a dangerous issue, because we don't really know how our brains work. If we set legal restrictions over this and then it turns out our brains work in a similar manner, then what?

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#442

Earlier quoted context omitted.

But they are not original works, they are wholly derived works of the training data set. Take that data set away and the algorithm is unable to produce a single original pixel. The fact that the derivation involves millions of works as opposed to a single one is immaterial for the copyright issue.

how is that any different from new human artist that study other artists work to learn a style or technique. In fact it used to be that the preferred way for painters to learn was to repeatedly copy paintings of masters.

What you and many other in the thread seem to be oblivious about is that algorithms are not people. Yes, it may come as a shock to autistic engineers, but the fact that a machine can do something to what a person does does not warant it equal protection under the law.

Copyright, and laws in general, exists to protect the human members of society not some abstract representation of them.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#443
post #405

Earlier quoted context omitted.

But they are not original works, they are wholly derived works of the training data set. Take that data set away and the algorithm is unable to produce a single original pixel. The fact that the derivation involves millions of works as opposed to a single one is immaterial for the copyright issue.

This argument's pedantic and problematic for artists; take away a human's "dataset" and processes and they are also unable to produce a single original "pixel".

[flagged]

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#444
post #413

Earlier quoted context omitted.

You took the images, encoded them in a computer process, and the result is able to reproduce some of those images. I fail to see why the size of the training set in bytes and the size of the model in bytes matters. Especially if, as other commenters have noted, much if the training data is repeated(mentions of thousands of mina Lisa's) so a straight division(training size/parameters size) says nothing about the bytes…

Except that you can't recreate them. At least not without a process that would be similar to asking an artist to create a replica of a painting. Just because photoshop has the right color palet available to recreate art, it doesn't mean the software itself is one big massive copyright infrigement against every art piece that exist.

Past a certain level of overfitting you can definitely recreate them just by asking for them by name. And it's possible to unintentionally or even intentionally overfit.

So it would be quite easy to make a trademark laundering operation, in theory.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#445

If they're so confident that copyright wouldn't apply to them, they should train their models using Disney content and set a proper precedent instead of using content from smaller artists.

Mickey's in there.

There you go. Try and generate a image of Mickey Mouse with his mouth closed using Stable Diffusion and you will see a verbatim image of him copied straight from Disney.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#446
post #2

“Sta­ble Dif­fu­sion con­tains unau­tho­rized copies of mil­lions—and pos­si­bly bil­lions—of copy­righted images.” That’s going to be hard to argue. Where are the copies? “Hav­ing copied the five bil­lion images—with­out the con­sent of the orig­i­nal artists—Sta­ble Dif­fu­sion relies on a math­e­mat­i­cal process called dif­fu­sion to store com­pressed copies of these train­ing images, which in turn are recom­bine…

> That’s going to be hard to argue. Where are the copies? In fairness, Diffusion is arguably a very complex entropy coding similar to Arithmetic/Huffman coding. Given that copyright is protectable even on compressed/encrypted files, it seems fair that the “container of compressed bytes” (in this case the Diffusion model) does “contain” the original images no differently than a compressed folder of images contains the…

> In fairness, Diffusion is arguably a very complex entropy coding similar to Arithmetic/Huffman coding.

Not the way it's used in Stable Diffusion models. Compressed data can be decompressed knowing only the decompression algorithm. To recover data from a stable diffusion model, you need to know the algorithm and the prompt.

A critical part of the information _isn't_ in the data you decompress, it has to come from you. (And this isn't that relevant, but it would be lossy, perceptual compression like jpeg or mp3, not lossless compression like Huffman or Arithmetic coding.)

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#447

Earlier quoted context omitted.

Makes sense. Does that mean OpenAI could essentially help fund Stability AI’s defense then?

I think they should. A lot of law is about precedent. AI-focused companies should get together, form a group, and have real experts giving expert testimony under oath. I also think that the courts will form expert committees consisting of real experts like Bengio, LeCun, etc. Any grifters should be avoided, but I am not sure if the judges will understand the difference.

Stability AI already drew this line with Dance Diffusion (like SD but for musicians) [0] trained on public domain music and on copyrighted music only with the permission from musicians?

The fact that Stabiliity is now creating an opt-out for artists after lifting and training on copyrighted / watermarked art without permission and creating a paid SaaS solution out of it, shows that not only they willfully trampled on the copyright of artists, but they have set themselves on a weak explanation on the 'fair use, transformative' argument since the LAION-5B model contains the copyrighted images which can output verbatim / highly similar digital art by the model.

The input from 'real experts like Bengio, LeCun' add little to no value in the case as digital art generated by a non-human is uncopyrightable and is public domain by default. [1] What sets the precedent is whether if using copyrighted content in a training set without permission from the author and outputting verbatim or highly similar derived works and commercializing that is fair use and not infringing.

If SD drew this line or musicians, then that should be the line drawn for digital art and code, and both Copilot, Midjourney, SD and DALL-E should be trained on public domain content or content with the permission from the content author or licenses that allow AI training.

So far, the 'real' grifters are Stability AI, OpenAI and Midjourney.

[0] https://techcrunch.com/2022/10/07/ai-music-generator-dance-d...

[1] https://www.copyright.gov/rulings-filings/review-board/docs/...

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#448

Earlier quoted context omitted.

Some years ago I had an idea to have a method of file sharing with strong plausible deniability from the sharer. The idea, in stage one, was to split a file into chunks and xor those with other random chunks (equivalent to a one-time pad), those chunks as well as the created random chunks then got shared around the networks, with nobody hosting both parts of a pair. The next stage is that future files inserted into t…

I think this touches on the core mismatch between the legal perspective and technical perspective. Yes, on a technical level, those chunks are random data. On the legal side, however, those chunks are illegal copyright infringement because that is their intent, and there is a process that allows the intent to happen. I can't really say it better than this post does, so I highly recommend reading it: https://ansuz.soo…

That's an interesting essay and I agree it goes to the heart of the question. There's clearly an interesting question, even in the colour domain: is someone infringing copyright if the data they themselves are sharing has a perfectly legitimate colour that is the basis of their sharing? That's the plausible deniability bit that's so important: "Yes your honour, I did share that chunk of random data, but I did so because it's part of this totally legitimately coloured file I was wanting to share. I had no idea that someone added a new colour to the block. Obviously, I'm only sharing the original colour block; prove otherwise". At some point, the court has to decide the colour of the block from the perspective of the accused, which allows a basis for deniability.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#449
post #273

Earlier quoted context omitted.

If I take a million copywritten images from magazines, cut them with scissors, and make a single collage, I would expect the resulting image to be fair use. Fair use is an affirmative defense, like self defense, where you justify your infringement. People are treating this like its a binary technical decision. Either it is or isn't a violation. Reality is that things are spectrums and judges judge. SD will likely be…

If I take a million copywritten images from magazines, cut them with scissors, and make a single collage, I would expect the resulting image to be fair use. That’s not how it works. Your collage would be fine if it was the only one since you used magazines you bought. Where you’d get into trouble is if you started printing copies of your collage and distributing them. In that case you’d be producing derived works and…

That’s not how fair use works. It’s not a binary switch where commercial derivatives automatically require licensing. Such a college would be ruled transformative and non competitive.

Me having bought the magazines also has nothing to do with anything. Would apply equally if they were gifted or free or stolen.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#450

Earlier quoted context omitted.

Great. Now the defence shows an artist that can recreate an image. Cool, now people who look at images get copyright suits filed against them for encoding those images in their heads.

Don't think stable Diffusion can reproduce any single image its trained on, not matter what prompts you use. It does have Mona lisa because of over fitting. But that's because there is too much Mona lisa on internet. These artist taking part in suit won't be able to recreat any of their work.

Does SD have to recreate the entire image for it to violate copyright?

As a thought experiment, imagine a variant of something like SD was used for music generation rather than images. It was trained on all music on spotify and it is marketed as a paid tool for producers and artists. If the model reproduces specific sounds from certain songs, e.g. the specific beat from a song, hook, or melody, it would seem pretty straightforward that the generated content was derivative, even though only a feature of it was precisely reproduced. I could be wrong but as far as i am aware you need to get permission to use samples. Even if the content is not published those sounds are being sold by the company as inspiration, and therefore that should violate copyright. The training data is paramount because if you trained the model on stuff you generated yourself or on stuff with appropriate CC license, the resulting work would not violate copyright, or you could at least argue independent creation.

In the feature space of images and art, SD is doing something very similar, so i can see the argument that it violates copyright even without reproducing the whole training data.

Overall, i think we will ultimately need to decide how we want these technologies used, what restrictions should be on the training data, etc, and then create new laws specifically for the new technology, rather than trying to shoehorn it into existing copyright law.

Post reply on HN