Live data from Hacker News

We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

stablediffusionlitigation.com

311–320 of 473 posts

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#311

“Hav­ing copied the five bil­lion images—with­out the con­sent of the orig­i­nal artists—Sta­ble Dif­fu­sion relies on a math­e­mat­i­cal process called dif­fu­sion to store com­pressed copies of these train­ing images, which in turn are recom­bined to derive other images.” This seems like it’s not an accurate description of what diffusion is doing. A diffusion model is not the same as compression. They’re implying t…

If you can put a bunch of large things together into a small file and then later (lossily) extract the large thing out of that smaller file, I'd argue that's compression, yeah. It doesn't really matter if it was intended to be art up as a compression algorithm or not in my opinion. If anything, this approach can be considered a revolution in lossy image compression, even though there's no real market for that at the moment.

If someone finds a way to reverse a hash, I'd also argue that hashing has now become a form of compression.

I think in 5 billion images there are more than enough common image areas to allow for average compression to become lower than a single byte. This is a lossy process, it does not need a complete copy of the source data, similar to how an MP3 doesn't contain most of the audio data fed into it.

I think the argument that SD revolves around lossless compression is quite an interesting one, even if the original code authors didn't realise that's what they were doing. It's the first good technical argument I've heard, at least.

All of those could've been prevented if the model was trained on public domain images instead of random people's copyrighted work. Even if this lawsuit succeeds, I don't think image generation algorithms will be banned. Some AI companies will just have spent a shitton of cash failing to get away with copyright violation, but the technology can still work for art that's either unlicensed or licensed in such a way that AI models can be trained based on it.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#312

This just wreaks of the Luddism of the 19th century all over again. https://en.wikipedia.org/wiki/Luddite Good luck stopping the inertia of progress.

I hope not. The Luddites were defeated by killing them.

> At one time there were more British soldiers fighting the Luddites than there were fighting Napoleon on the Iberian Peninsula.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#313

Does anyone else think it's grifty for a company to scrape up your (and other's) intellectual property, reconfigure it, and then attempt to sell it back to you for just $9.99 via dreambooth?

Depends on if “attempt to sell it back to you” is an accurate framing

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#314
post #88

Earlier quoted context omitted.

That's a huge understatement. 5 billion images to a model of 5GB. 1 byte per image. Let's see if one byte per image would constitute a copyright violation in other fields than neural networks.

Another thing worth referencing in this context might be hashing. If a few bytes per image are copyright infringement, then likely so is publishing checksums.

Once you start recreating copyrighted works from hashes this analogy becomes relevant, until then how can you compare the two when the distinguishing feature is its ability to reproduce the training data.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#315

Earlier quoted context omitted.

The pedantry gets tiring. If the AI can't recreate it exactly, it can recreate a likeness that is compelling enough that the average person would think it was the same. If it can't now, it will as it gets better. That's the point of using the training data.

Why does this argument apply to an Artificial Intelligence, but not a human one? A human is not breaking copyright just by being able recreate a copyrighted work they've studied.

It depends to what degree it's literal copying. See e.g. the Obama "Hope" poster. [1] Though that case is muddied by the fact that the artist lied about the source of his inspiration. Had it in fact been an older photo of JFK in a similar pose, there probably wouldn't have been a controversy.

[1] https://en.wikipedia.org/wiki/Barack_Obama_%22Hope%22_poster

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#316
post #44

Earlier quoted context omitted.

You could make the same argument that as long as you are using lossy compression you are unable to infringe on copyright.

if it's sufficiently lossy, yeah. don't know where you draw the line tho. maybe similar to fair use video clips.

Citing fair use is putting the cart before the horse here. The debate is around whether or not the stable diffusion training and generation processes can be considered transforming the worka to create a new one in the same way we do for humans which allows for the fair use of video clips. To say that it would be similar to fair use is assuming the outcome as evidence, aka begging the question.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#317
post #111

Earlier quoted context omitted.

why should a new right (the right to study the works) be granted without some compensation given back to society? The existing set of rights granted under copyright does not include this.

Yes, copyright protects expression, not ideas. Ideas are protected with patents.

It's implementations (supposedly) that are covered by patents. Ideas are not.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#318

“It is a par­a­site that, if allowed to pro­lif­er­ate, will make artists extinct.” This is the fundamentally flawed and misguided argument that can literally be applied to any technological progress to curtail advancement. Imagine if the medical tricorder (a device from Star Trek that does maybe 99% of what modern doctors do) is suddenly invented today. Doctors could use this argument to defend their livelihoods, bu…

[deleted]

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#319

It seems to me the communal voice of HN varies widely on copyright issues depending on who is getting sued and who is getting potentially hurt by violations. People who generally make less money than programmers - writers, artists, musicians - should stop their whining and their unfair uses of copyright to control their creative output. Programmers who are getting shafted by big corporations using their code to build…

I think the pushback from programmers has a different motivation. Programmers love to share code, but they don't want to share it with corporations who don't give back. We invented copyleft as a way to (ab)use the legal system to open up everything. We hate copyright and "love" "copyleft" as a means to weaken copyright. It would be like if artists gave away all of their art, except not to corporations who hog their c…

I'd argue that many artists are also fine with their art being reused and reposted elsewhere, as long as it's done with the necessary attribution.

For example, I don't think Sarah Ander­sen, one of the plaintiffs here, would've reached the popularity she's gained now if it wasn't for her comics being shared on meme sites and social media, for example, and I don't see any "do not repost" watermarks on her recent work like others that do object to resharing on other platforms have started doing.

I think many artists and copyleft programmers have the same positions. The biggest difference is that drawings are considered "art" and code is generally not, despite that fact a technical drawing and business logic are both hardly artistic and mostly an expression of skill whereas the demo scene, the indie gaming scene, and many online artists are very much about expressing themselves within a given set of boundaries.

When Microsoft steals code licensed to Github, people considered that to be a license dispute more than a copyright dispute. I'd argue there is no difference at all between artists and programmers when it comes to their work being absorbed and then reproduced by an AI company.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#320

Earlier quoted context omitted.

> It is, in short, a 21st-cen­tury col­lage tool. Interesting that they mention collages. IANAL but it was my impression that collages are derivative work if they incorporate many different pieces and only small parts of the original. Their compression argument seems more convincing.

Compression down to two bytes per image? You run into the pigeonhole argument. That level of compression can only work if there are less than seventy thousand different images in existence, total. Certainly there’s a deep theoretical equivalent between intelligence and compression, but this scenario isn’t what anyone means by “compression” normally.

When gzip turns my 10k character ASCII text file into a a 2kb archive, has it "compressed each character down to a fifth of a byte per character"? No, thats a misunderstanding of compression.

Just like gzip, training stable diffusion certainly removes a lot of data, but without understanding the effect of that transformation of the entropy of the data it's meaningless to say thing like "two bytes per image" because(like gzip) you need the whole encoded dataset to recover the image.

It's compressing many images into 10GB of data, not a single image into two bytes. This is directly analogous to what people usually mean by "compression"

Post reply on HN