Live data from Hacker News

We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

stablediffusionlitigation.com

151–160 of 473 posts

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#151

> Sta­ble Dif­fu­sion relies on a math­e­mat­i­cal process called dif­fu­sion to store com­pressed copies of these train­ing images, which in turn are recom­bined to derive other images. It is, in short, a 21st-cen­tury col­lage tool. Just no, that's not how any of that works. I guess that lie is convenient to legitimate the lawsuit.

It's a pretty funny assertion. The whole point of ML models is to take training data and learn something general from it, the common threads, such that it can identify/generate more things like the training examples. If the model were, as they assert, just compressing and reproducing/collaging training images then that would just indicate that the engineers of the model failed to prevent overfitting. So basically the…

As a side discussion, is there any research model which tries to do what they describe? Like overfitting to the maximum possible to create a way to compress data. It might be useful in different ways.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#152
post #20

Earlier quoted context omitted.

I don't think you have to reproduce an entire original work to demonstrate copyright violation. Think about sampling in hip hop for example. A 2 second sample, distorted, re-pitched, etc. can be grounds for a copyright violation.

Perhaps different media has different rules? You can’t necessarily apply music sampling rules to text, for example. Eg I don’t think incorporating a phrase from someone else’s poem into my poem would be grounds for a copyright violation.

"Copyright currently protects poetry just like it protects any other kind of writing or work of authorship. Poetry, therefore, is subject to the same minimal standards for originality that are used for other written works, and the same tests determine whether copyright infringement has occurred." [1]

[1] https://scholarship.law.vanderbilt.edu/vlr/vol58/iss3/13/

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#153
post #102

Earlier quoted context omitted.

> What else would they be trained on? why does it matter how it was trained? The question is, does the generative AI _output_ copyrighted images? Training is not a right that the copyright holder owns exclusively. Reproducing the works _is_, but if the AI only reproduces a style, but not a copy, then it isn't breaking any copyright.

Yes, because real artists are also allowed to learn from other paintings. No problem there, unless they recreate the exact work of others.

Banning AI from training on copyrighted works is also problematic because copyright doesn't protect ideas, it only protects expression. So the model has legitimate right to learn ideas (minus expression) from any source.

For example facts in the phonebook are not copyrighted, the authors have to mix fake data to be able claim copyright infringement. Maybe the models could finally learn how many fingers to draw on a hand.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#154

> Sta­ble Dif­fu­sion relies on a math­e­mat­i­cal process called dif­fu­sion to store com­pressed copies of these train­ing images, which in turn are recom­bined to derive other images. It is, in short, a 21st-cen­tury col­lage tool. Just no, that's not how any of that works. I guess that lie is convenient to legitimate the lawsuit.

How does it work then? :)

Diffusion models learn a transformation operator. The parameters are adjusted such that the operator maximises the evidence lower bound, or in other words, increasing the likelihood of observing a slightly less noisy version of the input.

The guidance component is a vector representation of the text that changes where we are in the sample space. A change in the sample space changes likelihood so for the different prompts the likelihood of the same output image for the same input image will be different.

Since the model is trained to maximise the ELBO, it will produce a change closer to the prompt.

A good way to think about it is this: given a classifier, I can select a target class and compute the derivative of the input with respect to the target class, and apply the derivative to the input. This puts it closer to my target class.

From the perspective of some models (score models), they produce a derivative of the density (of the samples), so it’s a bit similar to computing a derivative via classifier.

The above was concerned with what the NN was doing.

The algorithm applies the operator a number of steps, and progressively improves the image. In some probabilistic models, you can think of this as an inverse of stochastic gradient descent procedure (meaning a series of steps) that, with some stochasticity, reach a high value region (the density).

However, it turns out that learning this operation doesn’t have to be grounded in probability theory and graphical models.

As long as the NN learns a sufficiently good recovery operator, diffusion will construct something based on the properties of the dataset that has been used.

At no point however are there condensed representations of images since the NN is not learning to produce an image from zero in one step. It merely learns to recover some operation applied to the input.

For the probabilistic view, read Denoising Diffusion Probabilistic Networks and references, in particular langevin dynamics. It includes citations to score models as well.

For the non probabilistic component, read Cold diffusion.

For using the classifier gradient to update an image towards another class, read about adversarial generation via input gradients.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#155
post #26

Earlier quoted context omitted.

The difference here is that the images aren't stored, but rather an extremely abstract description of the image was used to very slightly adjust a network of millions of nodes in a tiny direction. No semblance of the original image even remotely exists in the model.

This is very much a 'color of your bits' topic, but I'm not sure why the internal representation matters. It's pretty trivial to recreate famous works like the Mona Lisa or Starry Night or Monet's Water Lily Pond. Obviously some representation of the originals exist inside the model+prompt. Why wouldn't that apply to other images in the training sets?

Because you're silently invoking additional data (the prompt + noise seed), which is not present in the training weights. You have the prompt + noise seed for any given output.

An MPEG codec doesn't contain every movie in the world just because it could represent them if given the right file.

The white light coming off a blank canvas also doesn't contain a copy of the Mona Lisa which will be revealed once someone obscures some of the light.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#156
post #98
post #72

Earlier quoted context omitted.

> This is the fundamentally flawed and misguided argument that can literally be applied to any technological progress to curtail advancement. No, this is the only fundamentally correct way to view this. Before the existence of the printing press, we didn't need copyright law. Yet all that the printing press did was make transcribing books by hand faster. Quantitative changes enabled by technology are qualitative chan…

name one qualitative change resulting from quantitative change that did not benefit the world as a whole?

All of the good changes have also come with new laws that forbid many of the bad uses thereof. Not every form of use of a technological invention is a net positive, and laws reflect that, by forbidding the negative uses.

The automobile revolutionized transportation, but also came with licensing requirements. (And more recently, we are finding to be responsible for a health and climate catastrophe, necessitating new restrictions on fuel economy, leaded gasoline, ICEs, etc.) You didn't need a license to walk or ride a bicycle, or ride a horse, but when we started putting people behind thousands of pounds of steel, all of a sudden we needed to come up with a myriad of new rules and restrictions on how automobiles could be used.

The printing press came with copyright laws. New and more destructive weapons and tools and chemicals came with more restrictions regarding their possession and expected use. The telephone and the computer combined allow robo calling and spam on an industrial level, and those particular uses of those new technologies are forbidden. Radio revolutionized communication, but we don't just let any random asshole blast static into the spectrum. We have narrowly curtailed, permitted and forbidden uses of it.

It would be far easier to name the technologies that net-benefited society, and did not need new rules around them, to prevent their destructive and damaging uses.

This one isn't looking to be one of them.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#157

> Sta­ble Dif­fu­sion relies on a math­e­mat­i­cal process called dif­fu­sion to store com­pressed copies of these train­ing images, which in turn are recom­bined to derive other images. It is, in short, a 21st-cen­tury col­lage tool. Just no, that's not how any of that works. I guess that lie is convenient to legitimate the lawsuit.

That's a lie, sure, but if they had instead claimed: The output of stable diffusion isn't possible without first examining millions of copyrighted images Then the suit looks a little more solid, because (as you pointed out) it isn't possible for the stable diffusion owner to know which of those copyright images had clauses that prevents stable diffusion trading and similar usage. The whole problem goes away once arti…

> The whole problem goes away once artists and photographers starting using a license that explicitly removes any use of the work as training data for any automated training.

A license which should be opt-in, not opt-out.

Of course, it’s opt-out because they know, fundamentally, that most artists would not want to opt-in.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#158
post #60

Earlier quoted context omitted.

> That’s going to be hard to argue. Where are the copies? In fairness, Diffusion is arguably a very complex entropy coding similar to Arithmetic/Huffman coding. Given that copyright is protectable even on compressed/encrypted files, it seems fair that the “container of compressed bytes” (in this case the Diffusion model) does “contain” the original images no differently than a compressed folder of images contains the…

And how that's different from gzip or base64, which can re-create original image when given appropriate input?

That’s my point, Diffusion[1] does seem to be “just like” gzip or base64.

And it would be illegal for me to sell or distribute zipped copies of images without the copyright holder’s consent. Similarly there might be an argument for why Diffusion[1] specifically can’t be built with copyrighted images.

[1] which is just one part of something like Stable Diffusion

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#159
post #21

Earlier quoted context omitted.

It doesn't matter if they exist as exact copies in my opinion. The law doesn't recognize a mathematical computer transformation as creating a new work with original copyright. If you give me an image, and I encrypt it with a randomly generated password, and then don't write down the password anywhere, the resulting file will be indistinguishable from random noise. No one can possibly derive the original image from it…

So what happens if you put a painting into a mechanical grinder? Is the shapeless pile of dust still copyrighted work? I don’t think so.

Maybe?

If you take a bad paper shredder that, say, shreds a photo into large re-usable chunks, run the photo through that, and tape the large re-usable chunks back together, you have a photo with the same copyright as before.

If you tape them together in a new creative arrangement, you might apply enough human creativity to create a new copyrighted work.

If you grind the original to dust, and then have a mechanical process somehow mechanically re-arrange the pieces back into an image without applying creativity, then the new mechanically created arrangement would, I suspect, be a derived work.

Of course, such a process don't really exist, so for the "shapeless dust" question, it's pretty pointless to think about. However, stable diffusion is grinding images down into neural networks, and then without a significant amount of human creativity involved, creating images reconstituted from that dust.

Perhaps the prompt counts as human creativity, but that seems fairly unlikely. After all, you can give it a prompt of 'dog' and get reconstituted dust, that hardly seems like it clears a bar.

Perhaps the training process somehow injected human creativity, but that also seems difficult to argue, it's an algorithm.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#160
post #146

Earlier quoted context omitted.

The training data set is indeed mandatory but that doesn't make the resulting model a derivative in itself. In fact the training is specifically made to remove derivatives.

Go to stablediffusionweb.com and enter "a person like biden" into the box. You will see a picture exactly like President Biden. That picture will have been derived from the trained images of Joe Biden. That cannot be in dispute.

Just because it generates you an image like Biden still does not make it a derivative either.

You can draw Biden yourself if you're talented and it's not considered a derivative of anything.

Post reply on HN