Earlier quoted context omitted.
Models for Stable Diffusion are about 2-8GB in size. 5 billion images means that every image gets about 1 byte. It seems to me that they're claiming here that Stability has somehow manage to store copies of these images in about 1 byte of space each. That's an incredible compression ratio!
It is a form compression that loses some much of the uniqueness which gives it the high ratio. If the concept is a little hard to grasp consider an AI model like a finite state machine, but it stores affinity and weights of the data's relationship to each other too. In GPT this is words and phrases, e.g. "Frodo Baggins" high affinity, "Frodo Superman" will be negligible. Now consider all words that may link to those…
We’ve filed a lawsuit challenging Stable Diffusion
441–450 of 473 posts
Re: We’ve filed a lawsuit challenging Stable Diffusion
#442Earlier quoted context omitted.
But they are not original works, they are wholly derived works of the training data set. Take that data set away and the algorithm is unable to produce a single original pixel. The fact that the derivation involves millions of works as opposed to a single one is immaterial for the copyright issue.
how is that any different from new human artist that study other artists work to learn a style or technique. In fact it used to be that the preferred way for painters to learn was to repeatedly copy paintings of masters.
Copyright, and laws in general, exists to protect the human members of society not some abstract representation of them.
Re: We’ve filed a lawsuit challenging Stable Diffusion
#443Earlier quoted context omitted.
But they are not original works, they are wholly derived works of the training data set. Take that data set away and the algorithm is unable to produce a single original pixel. The fact that the derivation involves millions of works as opposed to a single one is immaterial for the copyright issue.
This argument's pedantic and problematic for artists; take away a human's "dataset" and processes and they are also unable to produce a single original "pixel".
Re: We’ve filed a lawsuit challenging Stable Diffusion
#444Earlier quoted context omitted.
You took the images, encoded them in a computer process, and the result is able to reproduce some of those images. I fail to see why the size of the training set in bytes and the size of the model in bytes matters. Especially if, as other commenters have noted, much if the training data is repeated(mentions of thousands of mina Lisa's) so a straight division(training size/parameters size) says nothing about the bytes…
Except that you can't recreate them. At least not without a process that would be similar to asking an artist to create a replica of a painting. Just because photoshop has the right color palet available to recreate art, it doesn't mean the software itself is one big massive copyright infrigement against every art piece that exist.
So it would be quite easy to make a trademark laundering operation, in theory.
Re: We’ve filed a lawsuit challenging Stable Diffusion
#445If they're so confident that copyright wouldn't apply to them, they should train their models using Disney content and set a proper precedent instead of using content from smaller artists.
Mickey's in there.
Re: We’ve filed a lawsuit challenging Stable Diffusion
#446“Stable Diffusion contains unauthorized copies of millions—and possibly billions—of copyrighted images.” That’s going to be hard to argue. Where are the copies? “Having copied the five billion images—without the consent of the original artists—Stable Diffusion relies on a mathematical process called diffusion to store compressed copies of these training images, which in turn are recombine…
> That’s going to be hard to argue. Where are the copies? In fairness, Diffusion is arguably a very complex entropy coding similar to Arithmetic/Huffman coding. Given that copyright is protectable even on compressed/encrypted files, it seems fair that the “container of compressed bytes” (in this case the Diffusion model) does “contain” the original images no differently than a compressed folder of images contains the…
Not the way it's used in Stable Diffusion models. Compressed data can be decompressed knowing only the decompression algorithm. To recover data from a stable diffusion model, you need to know the algorithm and the prompt.
A critical part of the information _isn't_ in the data you decompress, it has to come from you. (And this isn't that relevant, but it would be lossy, perceptual compression like jpeg or mp3, not lossless compression like Huffman or Arithmetic coding.)
Re: We’ve filed a lawsuit challenging Stable Diffusion
#447Earlier quoted context omitted.
Makes sense. Does that mean OpenAI could essentially help fund Stability AI’s defense then?
I think they should. A lot of law is about precedent. AI-focused companies should get together, form a group, and have real experts giving expert testimony under oath. I also think that the courts will form expert committees consisting of real experts like Bengio, LeCun, etc. Any grifters should be avoided, but I am not sure if the judges will understand the difference.
The fact that Stabiliity is now creating an opt-out for artists after lifting and training on copyrighted / watermarked art without permission and creating a paid SaaS solution out of it, shows that not only they willfully trampled on the copyright of artists, but they have set themselves on a weak explanation on the 'fair use, transformative' argument since the LAION-5B model contains the copyrighted images which can output verbatim / highly similar digital art by the model.
The input from 'real experts like Bengio, LeCun' add little to no value in the case as digital art generated by a non-human is uncopyrightable and is public domain by default. [1] What sets the precedent is whether if using copyrighted content in a training set without permission from the author and outputting verbatim or highly similar derived works and commercializing that is fair use and not infringing.
If SD drew this line or musicians, then that should be the line drawn for digital art and code, and both Copilot, Midjourney, SD and DALL-E should be trained on public domain content or content with the permission from the content author or licenses that allow AI training.
So far, the 'real' grifters are Stability AI, OpenAI and Midjourney.
[0] https://techcrunch.com/2022/10/07/ai-music-generator-dance-d...
[1] https://www.copyright.gov/rulings-filings/review-board/docs/...
Re: We’ve filed a lawsuit challenging Stable Diffusion
#448Earlier quoted context omitted.
Some years ago I had an idea to have a method of file sharing with strong plausible deniability from the sharer. The idea, in stage one, was to split a file into chunks and xor those with other random chunks (equivalent to a one-time pad), those chunks as well as the created random chunks then got shared around the networks, with nobody hosting both parts of a pair. The next stage is that future files inserted into t…
I think this touches on the core mismatch between the legal perspective and technical perspective. Yes, on a technical level, those chunks are random data. On the legal side, however, those chunks are illegal copyright infringement because that is their intent, and there is a process that allows the intent to happen. I can't really say it better than this post does, so I highly recommend reading it: https://ansuz.soo…
Re: We’ve filed a lawsuit challenging Stable Diffusion
#449Earlier quoted context omitted.
If I take a million copywritten images from magazines, cut them with scissors, and make a single collage, I would expect the resulting image to be fair use. Fair use is an affirmative defense, like self defense, where you justify your infringement. People are treating this like its a binary technical decision. Either it is or isn't a violation. Reality is that things are spectrums and judges judge. SD will likely be…
If I take a million copywritten images from magazines, cut them with scissors, and make a single collage, I would expect the resulting image to be fair use. That’s not how it works. Your collage would be fine if it was the only one since you used magazines you bought. Where you’d get into trouble is if you started printing copies of your collage and distributing them. In that case you’d be producing derived works and…
Me having bought the magazines also has nothing to do with anything. Would apply equally if they were gifted or free or stolen.
Re: We’ve filed a lawsuit challenging Stable Diffusion
#450Earlier quoted context omitted.
Great. Now the defence shows an artist that can recreate an image. Cool, now people who look at images get copyright suits filed against them for encoding those images in their heads.
Don't think stable Diffusion can reproduce any single image its trained on, not matter what prompts you use. It does have Mona lisa because of over fitting. But that's because there is too much Mona lisa on internet. These artist taking part in suit won't be able to recreat any of their work.
As a thought experiment, imagine a variant of something like SD was used for music generation rather than images. It was trained on all music on spotify and it is marketed as a paid tool for producers and artists. If the model reproduces specific sounds from certain songs, e.g. the specific beat from a song, hook, or melody, it would seem pretty straightforward that the generated content was derivative, even though only a feature of it was precisely reproduced. I could be wrong but as far as i am aware you need to get permission to use samples. Even if the content is not published those sounds are being sold by the company as inspiration, and therefore that should violate copyright. The training data is paramount because if you trained the model on stuff you generated yourself or on stuff with appropriate CC license, the resulting work would not violate copyright, or you could at least argue independent creation.
In the feature space of images and art, SD is doing something very similar, so i can see the argument that it violates copyright even without reproducing the whole training data.
Overall, i think we will ultimately need to decide how we want these technologies used, what restrictions should be on the training data, etc, and then create new laws specifically for the new technology, rather than trying to shoehorn it into existing copyright law.