Live data from Hacker News

We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

stablediffusionlitigation.com

321–330 of 473 posts

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#321
post #254

“It is a par­a­site that, if allowed to pro­lif­er­ate, will make artists extinct.” This is the fundamentally flawed and misguided argument that can literally be applied to any technological progress to curtail advancement. Imagine if the medical tricorder (a device from Star Trek that does maybe 99% of what modern doctors do) is suddenly invented today. Doctors could use this argument to defend their livelihoods, bu…

> This is the fundamentally flawed and misguided argument that can literally be applied to any technological progress to curtail advancement Let's stop for a moment and define advancement (or "progress", as it's sometimes called). It's always tacit, and never explicitly defined, and I think it bears examination. By advancement/progress, I'm taking the argument to mean "betterment". i.e. When we say "advances in scien…

Maybe AI is the visual shortcut of an Excel Pivot Table: people use it, slice and dice, and get further insights to some purpose.

It's a tool. Folks get excited about statistics, massive datasets, and computer science is hip again.

Would we not want a push for folks to experience the exacting caress of an unforgiving compiler?

I thought this stuff would be easy!

Hopefully what doesn't happen is a fragmentation of folks into content caverns, where they may gaze into a mirror and see exactly what they wish, day after day. A literal instantiation of Plato's Caves, where scientific progress is frozen and forgotten.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#322

The discussion here has been of a somewhat casual nature: - expressing personal feelings about the lawsuit - honestly sharing one's lack of understanding of the legal or technical issues involved - making brief, unsubstantiated claims about the lawsuit's merits - discussing HN response to this story (as this comment) I was really looking forward to educating myself on this topic from the HN comments, but it appears I…

Probably because:

1. No one knows. Real expert IP lawyers I know tend to think it's probably OK with various caveats.

2.) Under even simpler scenarios fair use is a pretty fuzzy line even for things that are better understood and often end up litigated on a case basis. Someone else gave various examples from music. I cited the Obama "Hope" poster elsewhere. Very little is absolute at the margins.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#323
post #150
post #2

“Sta­ble Dif­fu­sion con­tains unau­tho­rized copies of mil­lions—and pos­si­bly bil­lions—of copy­righted images.” That’s going to be hard to argue. Where are the copies? “Hav­ing copied the five bil­lion images—with­out the con­sent of the orig­i­nal artists—Sta­ble Dif­fu­sion relies on a math­e­mat­i­cal process called dif­fu­sion to store com­pressed copies of these train­ing images, which in turn are recom­bine…

> That’s going to be hard to argue. Where are the copies? If you take that tack, I'll go one step further back in time and ask "Where is your agreement from the original author who owns the copyright that you could use this image in the way you did?" The fact that there is suddenly a new way to "use an image" (input to a computer algorithm) doesn't mean that copyright magically doesn't also apply to that usage. A can…

My assumption would be 'fair use'. Artists themselves make use of this extremely often, like when doing paintovers on copyrighted images (VERY common), fan art where they paint trademarked characters (also VERY common). The are often done for commission as well.

AFAIK, downloading and learning from images, even copyrighted images, fall under fair use, this is how practically every artist today learns how to draw.

Stable Diffusion does not create 1:1 copies of artwork it has been trained on, and its purpose is quite the opposite, there may be cases where the transformative aspect of a generated image may be argued as not being transformative enough, but so far I've only seen one such reproducable image, which would be the 'bloodborne box art' prompt, which was also mentioned in this discussion.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#324

Oddly, there's no mention of CLIP in the post and filing, where CLIPText is the real secret of Stable Diffusion's ability to correlate text to image for all the iterations targeted by the plantiffs. I suspect adding OpenAI as a defendant would make things a tad harder legally.

SD 2.x doesn't use CLIPText anymore, although if they don't know the difference between 1.x and 2.x their argument isn't going to be very good.

The lawsuit doesn't make that distinction and the DeviantArt model is almost certainly based on SD 1.x

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#325

Does anyone else think it's grifty for a company to scrape up your (and other's) intellectual property, reconfigure it, and then attempt to sell it back to you for just $9.99 via dreambooth?

Depends on if “attempt to sell it back to you” is an accurate framing

I think it is. Take away the training and there is no product.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#326

Earlier quoted context omitted.

lol thinking about this more: I understand people’s livelihoods are potentially at stake, but what a shame it would be if we find AGI, even consciousness but have to shut it down because of a copyright dispute.

I think the result will be image sharing websites where you have to agree to have your image read into the model. I think it is likely github will do the same with copilot.

Image sharing sites routinely steal artwork from the web. My business has a unique logo with the business name in it. It has repeatedly shown up on such sites, despite repeated DMCA takedown requests.

Simply appearing on a shared hosting site should not be enough.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#327

Earlier quoted context omitted.

I believe Copilot was giving exact copies of large parts open source projects, without the license. Are image generators giving exact (or very similar) copies of existing works? I feel like this is the main distinction.

Not large parts of open source projects. It was one function that was pretty well known and replicated. The author prompted with a part of the code, and the model finished the rest including the original comments. There are two issues here - the model needs to be carefully prompted (goaded) into copyright violation, so it is instigated to do it by excessive quoting from the original - the replicated codes are usually…

Copilot didn't just spit out the fast inverse square root, it spat out someone's entire "about" page in HTML, name and all. This was just some guy's blog, not a commonly replicated algorithm from a book.

Furthermore, copyright infringement doesn't stop being copyright infringement if you do it based on someone else's copyright infringement. Just become someone else decided to rip the contents of a CD and upload it to a website doesn't mean I'm now allowed to download it from that website again.

Copyright law does include an originality floor, you can't copyright a letter or a shape unless you're a billion dollar startup and in the same way that you can't copyright fizzbuzz or hello world. I don't think that's relevant for many algorithms Copilot will generate for you, though.

If simple work doesn't deserve protection, the pop music industry with their generic lyrics and simple tunes may be in big trouble. Disney as well, with their simplistic cartoon characters like Donald Duck and Mickey Mouse.

Personally, I think copyright laws are extremely damaging in their duration and restrictions. IP law in a small amount of countries actually allows for patenting algorithms, which is equally silly. International IP law currently gets in the way of society in my opinion.

However, without short term copyright neither programmers nor artists will be happy and I don't think anyone but knock-off companies will be happy with such an arrangement. Five or ten years is long enough for copyright in my book, but within those five or ten years copyright must remain protected.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#328

Earlier quoted context omitted.

> That’s going to be hard to argue. Where are the copies? In fairness, Diffusion is arguably a very complex entropy coding similar to Arithmetic/Huffman coding. Given that copyright is protectable even on compressed/encrypted files, it seems fair that the “container of compressed bytes” (in this case the Diffusion model) does “contain” the original images no differently than a compressed folder of images contains the…

Great. Now the defence shows an artist that can recreate an image. Cool, now people who look at images get copyright suits filed against them for encoding those images in their heads.

Don't think stable Diffusion can reproduce any single image its trained on, not matter what prompts you use.

It does have Mona lisa because of over fitting. But that's because there is too much Mona lisa on internet.

These artist taking part in suit won't be able to recreat any of their work.

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#329
post #2

“Sta­ble Dif­fu­sion con­tains unau­tho­rized copies of mil­lions—and pos­si­bly bil­lions—of copy­righted images.” That’s going to be hard to argue. Where are the copies? “Hav­ing copied the five bil­lion images—with­out the con­sent of the orig­i­nal artists—Sta­ble Dif­fu­sion relies on a math­e­mat­i­cal process called dif­fu­sion to store com­pressed copies of these train­ing images, which in turn are recom­bine…

> That’s going to be hard to argue. Where are the copies? In fairness, Diffusion is arguably a very complex entropy coding similar to Arithmetic/Huffman coding. Given that copyright is protectable even on compressed/encrypted files, it seems fair that the “container of compressed bytes” (in this case the Diffusion model) does “contain” the original images no differently than a compressed folder of images contains the…

There's a key difference. A compression algorithm is made to be reversible. The point of compressing an MP3 is to be able to decompress as much of the original audio signal as possible.

Stable Diffusion is not made to decompress the original and actually has no direct mechanism for decompressing any originals. The originals are not present. The only thing present is an embedding of key components of the original in a multi-dimensional latent space that also includes text.

This doesn't mean that the outputs of Stable Diffusion cannot be in violation of a copyright, it just means that the operator is going to have to direct the model towards a part of that text/image latent space that violates copyright in some manner... and that the operator of the model, when given an output that is in violation of copyright, is liable for publishing the image. Remember, it is not a violation of copyright to photocopy an image in your house... it's a violation when you publish that image!

Re: We’ve filed a law­suit chal­leng­ing Sta­ble Dif­fu­sion

#330

Earlier quoted context omitted.

The pedantry gets tiring. If the AI can't recreate it exactly, it can recreate a likeness that is compelling enough that the average person would think it was the same. If it can't now, it will as it gets better. That's the point of using the training data.

That is not the point of using the training data. It's specifically trained to not do that. See https://openai.com/blog/dall-e-2-pre-training-mitigations/ "Preventing Image Regurgitation".

That's probably a very relevant point. (I'm guessing.) If I ask for an image of a red dragon in the style of $ARTIST, and the algorithm goes off and says "Oh, I've got the perfect one already in my data"--or even "I've got a few like that, I'll just paste them together"--that's a problem.
Post reply on HN