Live data from Hacker News

We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

stablediffusionlitigation.com

141–150 of 473 posts

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#141

So what is the end goal of this? For copyright to transfer every step? That precedent happens, then what? Licensing schemes get set up and any piece of media that is put into these systems will result in the artist getting some kind of payment in return. Cool, that sound great. Except... who's paying? The conglomerates who already have a bunch of IP they can feed into those systems, who can afford to purchase or thro…

Mostly I agree, but:

> Copyright is a prison built for artists by big business, successfully marketed to artists as being a home.

I think (continuing this analogy) that copyright is indeed a home, but very few artists can afford to buy their own home, so they rent from corporate landlords, and the bigger ones are the worst ones to be tenants of.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#142

Earlier quoted context omitted.

lol thinking about this more: I understand people’s livelihoods are potentially at stake, but what a shame it would be if we find AGI, even consciousness but have to shut it down because of a copyright dispute.

what if they shut us down because of a copyright dispute? :-)

Seriously!!!

I didn’t say it cuz I didn’t think it would resonate, but it’s a whole new world we are quickly entering.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#143
post #21

Earlier quoted context omitted.

It doesn't matter if they exist as exact copies in my opinion. The law doesn't recognize a mathematical computer transformation as creating a new work with original copyright. If you give me an image, and I encrypt it with a randomly generated password, and then don't write down the password anywhere, the resulting file will be indistinguishable from random noise. No one can possibly derive the original image from it…

Some years ago I had an idea to have a method of file sharing with strong plausible deniability from the sharer. The idea, in stage one, was to split a file into chunks and xor those with other random chunks (equivalent to a one-time pad), those chunks as well as the created random chunks then got shared around the networks, with nobody hosting both parts of a pair. The next stage is that future files inserted into t…

I think this touches on the core mismatch between the legal perspective and technical perspective.

Yes, on a technical level, those chunks are random data. On the legal side, however, those chunks are illegal copyright infringement because that is their intent, and there is a process that allows the intent to happen.

I can't really say it better than this post does, so I highly recommend reading it: https://ansuz.sooke.bc.ca/entry/23

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#144
post #72

Earlier quoted context omitted.

> This is the fundamentally flawed and misguided argument that can literally be applied to any technological progress to curtail advancement. No, this is the only fundamentally correct way to view this. Before the existence of the printing press, we didn't need copyright law. Yet all that the printing press did was make transcribing books by hand faster. Quantitative changes enabled by technology are qualitative chan…

Artists do not have an inalienable right to be paid to do art. There are many reasons why you might argue that the tech is harmful, but that it "will make artists extinct" is not a good one.

And authors don't have any natural right to prevent people from copying their books.

And yet, we have decided that society is better off when authors can make money off their work.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#145

Sometimes I have to wonder about the hypocrisy you can see on HN threads. When its software development, many here seem to understand the merits of a similar lawsuit against Copilot[1], but as soon as its a different group such as artists, then it's "no, that's not how a NN works" or "the NN model works just the same way as a human would understand art and style." [1] https://news.ycombinator.com/item?id=34274326

I believe Copilot was giving exact copies of large parts open source projects, without the license. Are image generators giving exact (or very similar) copies of existing works? I feel like this is the main distinction.

Not large parts of open source projects. It was one function that was pretty well known and replicated. The author prompted with a part of the code, and the model finished the rest including the original comments.

There are two issues here

- the model needs to be carefully prompted (goaded) into copyright violation, so it is instigated to do it by excessive quoting from the original

- the replicated codes are usually boilerplate, common approaches or "famous" examples from books; in other words they are examples that appear in multiple places in the training set as opposed to just once

Do generic codes, boilerplate and API calls deserve protection? Maybe the famous examples do, but not every replicated code does.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#146

Earlier quoted context omitted.

But they are not original works, they are wholly derived works of the training data set. Take that data set away and the algorithm is unable to produce a single original pixel. The fact that the derivation involves millions of works as opposed to a single one is immaterial for the copyright issue.

The training data set is indeed mandatory but that doesn't make the resulting model a derivative in itself. In fact the training is specifically made to remove derivatives.

Go to stablediffusionweb.com and enter "a person like biden" into the box. You will see a picture exactly like President Biden. That picture will have been derived from the trained images of Joe Biden. That cannot be in dispute.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#147
post #88
post #44

Earlier quoted context omitted.

You could make the same argument that as long as you are using lossy compression you are unable to infringe on copyright.

That's a huge understatement. 5 billion images to a model of 5GB. 1 byte per image. Let's see if one byte per image would constitute a copyright violation in other fields than neural networks.

The distribution of the bytes matters a bit here. In theory the model could be over trained against one copyrighted work such that it is almost perfectly preserved within the model.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#148
post #2

“Sta­ble Dif­fu­sion con­tains unau­tho­rized copies of mil­lions—and pos­si­bly bil­lions—of copy­righted images.” That’s going to be hard to argue. Where are the copies? “Hav­ing copied the five bil­lion images—with­out the con­sent of the orig­i­nal artists—Sta­ble Dif­fu­sion relies on a math­e­mat­i­cal process called dif­fu­sion to store com­pressed copies of these train­ing images, which in turn are recom­bine…

You seem to be under the impression that SD can only generate original art. However, it will literally recreate existing paintings for you if you just prompt it with the title. Identical composition and everything.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#149

Earlier quoted context omitted.

Great. Now the defence shows an artist that can recreate an image. Cool, now people who look at images get copyright suits filed against them for encoding those images in their heads.

Just because I look at an image does not mean that I can recreate it. storing it in the training data means the AI can recreate it. There's a world of difference that you are just writing off.

No, it means there is a 512 bit number you can combine with the training data to reproduce a reasonable though not exact likeness (attempts to use SD and others as compression algorithms show they're pretty bad at it, because while they can get "similar" they'll outright confabulate details in a plausible looking way - i.e. redrawing the streets of San Francisco in images of the golden gate bridge).

Which of course then arrives at the problem: the original data plainly isn't stored in a byte exact form, and you can only recover it by providing an astounding specific input string (the 512 bit latent space vector). But that's not data which is contained within Stable Diffusion. It's equivalent to trying to sue a compression codec because a specific archive contains a copyrighted image.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#150
post #2

“Sta­ble Dif­fu­sion con­tains unau­tho­rized copies of mil­lions—and pos­si­bly bil­lions—of copy­righted images.” That’s going to be hard to argue. Where are the copies? “Hav­ing copied the five bil­lion images—with­out the con­sent of the orig­i­nal artists—Sta­ble Dif­fu­sion relies on a math­e­mat­i­cal process called dif­fu­sion to store com­pressed copies of these train­ing images, which in turn are recom­bine…

> That’s going to be hard to argue. Where are the copies?

If you take that tack, I'll go one step further back in time and ask "Where is your agreement from the original author who owns the copyright that you could use this image in the way you did?"

The fact that there is suddenly a new way to "use an image" (input to a computer algorithm) doesn't mean that copyright magically doesn't also apply to that usage.

A canonical example is the fact that television programs like "WKRP in Cincinnati" can't use the music licenses from the television broadcast if they want to distribute a DVD or streaming version--the music has to be re-licensed.

Post reply on HN