Live data from Hacker News

We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

stablediffusionlitigation.com

371–380 of 473 posts

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#371
post #2

“Sta­ble Dif­fu­sion con­tains unau­tho­rized copies of mil­lions—and pos­si­bly bil­lions—of copy­righted images.” That’s going to be hard to argue. Where are the copies? “Hav­ing copied the five bil­lion images—with­out the con­sent of the orig­i­nal artists—Sta­ble Dif­fu­sion relies on a math­e­mat­i­cal process called dif­fu­sion to store com­pressed copies of these train­ing images, which in turn are recom­bine…

Models for Stable Diffusion are about 2-8GB in size. 5 billion images means that every image gets about 1 byte.

It seems to me that they're claiming here that Stability has somehow manage to store copies of these images in about 1 byte of space each. That's an incredible compression ratio!

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#372

Earlier quoted context omitted.

My assumption would be 'fair use'. Artists themselves make use of this extremely often, like when doing paintovers on copyrighted images (VERY common), fan art where they paint trademarked characters (also VERY common). The are often done for commission as well. AFAIK, downloading and learning from images, even copyrighted images, fall under fair use, this is how practically every artist today learns how to draw. Sta…

> when doing paintovers on copyrighted images (VERY common) What are you talking about? I've been doing drawing and digital painting as a hobby for a long time and tracing is absolutely not "VERY common". I don't know anybody who has ever done this. > fan art where they paint trademarked characters (also VERY common) This is true in the sense that many artists do it (besides confusing trademark law and copyright law:…

>and tracing is absolutely not "VERY common"

Paintover does not have to mean actual 'tracing', a LOT of artists use photos as direct references and paint over them in a separate layer, keeping the composition, poses, colors very close to the original while still changing details and style enought to make it transformative enough to be considered a 'new work'.

Here are two examples of artist Sam Yang using two still frames from the tv show Squid Game and painting over those, the results which he then sells as prints:

https://www.inprnt.com/gallery/samdoesarts/the-alleyway/ https://www.inprnt.com/gallery/samdoesarts/067/

That said, you could even get away with less transformation and still have it be considered original work, take Andy Warhol's 'Orange Marilyn' and 'Portrait of Mao', those are inked and flat color changes over photographs.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#373
post #351

Earlier quoted context omitted.

I doubt Microsoft sees fragments of Windows source code as a particular crown jewel these days. That said, some of it is decades old code that was intended for the public to see (unlike, presumably, anything in a public GitHub repository). And some of it is presumably third-party code licensed to Microsoft that was likewise never intended for public viewing. So, while it would be a good gesture on the part of Microso…

Private third-party GitHub repos is another good example. If licenses don't apply to training data, as GitHub has asserted, why not use those too? Do they think they'll get in trouble over it? Why doesn't the same trouble apply to my publicly-readable GPL-licensed code?

I assume there's something in their terms of service about not poking around in private repos and using the code even for internal purposes except for necessary maintenance like backups, court orders, etc.

I am not a lawyer but I also assume Microsoft's position, at least in part, is that they can download and use code in GitHub public repos just like anyone else can and developing a public service based on training with that (and a lot of other) code isn't redistributing that code.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#374
post #362

So what is the end goal of this? For copyright to transfer every step? That precedent happens, then what? Licensing schemes get set up and any piece of media that is put into these systems will result in the artist getting some kind of payment in return. Cool, that sound great. Except... who's paying? The conglomerates who already have a bunch of IP they can feed into those systems, who can afford to purchase or thro…

You're arguing that artists have a shitty home, therefore it's not worth protecting as those AI companies are trying to take even that from them. And you're somehow trying to sound like you're pro artists in all this. Please, listen to yourself.

I wouldn't try reading too much into the pithy poetry I added at the last minute to make a broad, perhaps not particularly clear point about how copyright has been twisted to only serve established conglomerates rather than individuals.

My main beef with the approach being taken by a lot of artists towards AI art generators, using copyright in an attempt to kick the tools in the shin (notably not actually kill it, only perhaps slow it down a bit) could set legal precedent that would make it worse for individual artists and smaller groups by putting the most useful and powerful variations of the technology exclusively in the hands of intellectual property hoarders like Disney. As opposed to a more open approach where the possibility exists for useful generators to exist for free.

I'm not anti artist, I do genuinely think that this outcome would make things worse for them and better for the companies that already exploit them.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#375

Earlier quoted context omitted.

Great. Now the defence shows an artist that can recreate an image. Cool, now people who look at images get copyright suits filed against them for encoding those images in their heads.

Don't think stable Diffusion can reproduce any single image its trained on, not matter what prompts you use. It does have Mona lisa because of over fitting. But that's because there is too much Mona lisa on internet. These artist taking part in suit won't be able to recreat any of their work.

I think there's a chance they might be able to recreate some simpler work if they make the prompts specific enough. When you set up a prompt you're essentially telling the system what you want it to generate - if you prompt it with enough specificity you might be able to just recreate the image you had.

Kind of like recreating your image one object at a time. It might not be exact, but close enough.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#377
post #35
post #26

Earlier quoted context omitted.

The difference here is that the images aren't stored, but rather an extremely abstract description of the image was used to very slightly adjust a network of millions of nodes in a tiny direction. No semblance of the original image even remotely exists in the model.

there are some artists with very strong, recognizable styles. if you provide one of these artists' name in your prompt and get a result back that employs their strong, recognizable style, i think that demonstrates that the network has a latent representation of the artists work stored inside of it.

That seems to indicate to me that the original work is actually not under copyright, since if it is the only method of achieving such an image in such a style, then there is no originality to be copyrighted.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#378

Does anyone else think it's grifty for a company to scrape up your (and other's) intellectual property, reconfigure it, and then attempt to sell it back to you for just $9.99 via dreambooth?

If they also give you the means to just do it yourself?

Imo something like dall-e or midjourney is much worse.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#380
post #351

Earlier quoted context omitted.

I doubt Microsoft sees fragments of Windows source code as a particular crown jewel these days. That said, some of it is decades old code that was intended for the public to see (unlike, presumably, anything in a public GitHub repository). And some of it is presumably third-party code licensed to Microsoft that was likewise never intended for public viewing. So, while it would be a good gesture on the part of Microso…

Private third-party GitHub repos is another good example. If licenses don't apply to training data, as GitHub has asserted, why not use those too? Do they think they'll get in trouble over it? Why doesn't the same trouble apply to my publicly-readable GPL-licensed code?

Copyright is not the only law. Something might be permitted by copyright law (as fair use, an implied license, etc)-yet simultaneously violate other laws-breach of contract, misappropriation of trade secrets, etc.
Post reply on HN