Live data from Hacker News

Stable Diffusion is a big deal

simonwillison.net

171–180 of 488 posts

Re: Stable Diffusion is a big deal

#171

Earlier quoted context omitted.

There are really significant, novel copyright issues implicated by these large generative models trained on other people’s IP. If you take a step back, you can see that there are different ways to frame what is happening. One frame is: “Defendant built an algorithm that memorized features of Plaintiff’s IP. Defendant’s algorithm recombines parts of those features in order to produce works in the same domain that comp…

Artist today use the exact same method of learning from other peoples artwork to generate new artwork and styles. These models are learning just like any artist learns and then producing new content.

This is patently, obviously _wrong_ for anyone who has tried learning any artistic skill in their life. Sorry to be this straightforward, but it gets on my nerves every time I read it.

If you tried learning, let's say, the chiaroscuro technique from Caravaggio you'd be analyzing the way the painter simulated volumetric space by using white and dark tones in place of natural lighting and shadows. You wouldn't even think of splitting the whole painting into puzzle size pieces while checking how many how those look similar when put close one another.

Given somewhat decent painting skills, you'd be able to steadily apply this technique for the rest of your life just by looking at a very small sample of Caravaggio's corpus.

On the other hand if you tried removing even just a single work from the original Stable Diffusion data set you used to generate your painting, it would be absolutely impossible to recreate a similar enough picture even by starting from the same prompt and seed values.

Given how smart some of the people working on this are, I'm starting to believe they're intentionally playing dumb to make sure nobody is going ask them to prove this during a copyright infringement case.

Re: Stable Diffusion is a big deal

#172

Maybe this current explosion in the relevance and visibility of this kind of AI model will finally lead us to rethink how insanely nonsensical our IP systems are. I'm not holding my breath, but there's hope that this sort of thing will (combined with situations like the HBO debacle) clarify the need for massive IP reform in the cultural zeitgeist. The problem here isn't that the model was trained on copyrighted works…

Why do people keep bringing up copyright anyway? It seems pretty clear that the images being generated by StableDiffusion are transformative so it's protected under Fair Use. Am I missing something here?

Fair Use is a defense, not a right. Just being transformative isn't enough here, that's just one of many different factors that needs to be checked. Furthermore Fair Use includes evaluating the effect an infringement has upon original work's value.

So when your image generator keeps spitting out Gettyimages watermarks, while you are building a service that is in direct competition with Gettyimages for stock images, there is an argument to be made that Fair Use really doesn't apply here. As what you are doing is essentially stealing Gettyimages' work, AI laundering it and selling it back to their previous customers.

With StableDiffusion a Fair Use defense might have an easier time, as the results are released to the public. But it's still not exactly clear cut. If you type in "Mona Lisa", you'll still get something that looks like a copy of the Mona Lisa, not like an original work.

Re: Stable Diffusion is a big deal

#173

Why do Stable Diffusion images all sorta look the same. They look either deformed or like pastels . I tried it a few times on some free sites, and it was not accurate. I would put in 5 keywords and maybe it would get 2 of them right. These are common search terms.

You have to really spend a lot of time to build the prompt to get something nice and not generic.

I encourage you to check out the subreddit where people share some prompts that you can reuse.

There's usually only a part that describes what is in the image, while half the prompt will be used to describe how does the image look.

Thinking of this one for example: https://www.reddit.com/r/StableDiffusion/comments/wp26lp/i_s... I have reused it successfully.

Re: Stable Diffusion is a big deal

#174
post #10

Earlier quoted context omitted.

What about using this tech for ideation and artists for production? You could use Stable Diffusion et al to create new characters based on a prompt, then farm the concept out to artists to produce individual works. Kind of like hiring a super expensive agency to design your new logo or brand identity, then using a stable of in-house designers to translate the concept into UI, ads, etc.

This is useless, you still have humans in the loop.

Useless? One can spend days/weeks/months enumerating different concepts of a design due to a roundtrip between “maybe we try ” and an image. Now it’s literally minutes and a designer can draw you some ideas right at your office.

Re: Stable Diffusion is a big deal

#175
> Stable Diffusion has been trained on millions of copyrighted images scraped from the web.

My brain has been trained on even more copyrighted material. Every book I read, every tv show I watch, the toys I played with as a child. It's hard to imagine that I could come up with anything that is not inspired by copyrighted work.

Re: Stable Diffusion is a big deal

#176

Regarding the ethics. Thought experiment: I look at a copyrighted image, get inspiration, and then draw a new image (which is not the same as the copyrighted image). Is that ethical? But if I do this millions of times, after looking at millions of copyrighted images, is it still ethical (as long as I don't recreated copyrighted images)? Tough questions!

The ethics partly depends on the analogy you choose. I can write a book about a boy wizard's adventures wizard school and that's legal, but if I call them Harry Potter it isn't. I can create Harry Potter fanart and distribute it online pretty freely - but slap it on a mug and sell it, and that's illegal. I can record an audio description of a painting that's as detailed as I like and it's legal to distribute - but ta…

Harry Potter is a trademark. It protects and distinguishes identity and authority. It is much harder to get a trademark, and is also harder to fairly use one. Different issue.

If your fan art is infringing, it was infringing whether or not it was on a mug or on dropbox.

Photograph one is not true. For commentary it can be by audio, or printed on a mug or whatever, commentary is transformative. Go out and take a photo of the world outside, if you live in a city you've captured thousands of copyrighted materials in your image. Maybe it captures someones painting, maybe it doesn't. Whether its fair or not is if it's transformative, the format doesn't matter.

I'm not seeing an ethical difference from any of this. Or did I miss the point?

Re: Stable Diffusion is a big deal

#177
post #56

> No-one expected creative AIs to come for the artist jobs first, but here we are! Maybe that's because we never really thought about it. In hindsight, it's only logical. For an artistic rendering, correctness doesn't matter much nor does understanding the model. For flying a plane or driving a car or just transforming code from one language to another, it very much does.

Surely the best people to create new art from models like this would be artists themselves?

Wouldn't this create more jobs for artists? It's just a different tool

Re: Stable Diffusion is a big deal

#179
post #11

Earlier quoted context omitted.

For every tool, a better output will always be created by people who specialize in a craft. Just as photoshop revolutionized photography, to this day you can tell the difference easily between and bad 'shops. In video games upon release everyone is bad and its an even playing field. But as people practice and gain experience they improve their usage, refine their approaches. You eventually see metas develop and best…

Yes, but art school just turned into a one-semester course.

Not yet, but I can definitely imagine a future where these tools get more capable and refined, to the point where all the shortcomings listed above will be overcome. Knowledge about cameras and scene composition are already encoded in the networks to some degree, it just needs to become more accessible. There's probably also a better way to seed new images than by starting with random noise, so we could get similar variations easier. We have already made the big step towards creativity and real world understanding of objects and their lighting, the remaining issues are more technical and unless we are incredibly unlucky and run into a true show-stopper, we'll probably all have access to a high quality digital artist that can reduce production times dramatically.

Re: Stable Diffusion is a big deal

#180

I'm no artist.. though I'll admit i dabble, but is it just me or does everything that is generated in this seem to lack emotion ? I have seen some very impressive pictures but nothing that seems to "emote"

Yep it’s missing intent, just like machine translation. There somehow seems to be less information than there is in the input. I think more people will start to notice it as time goes by.

> There somehow seems to be less information than there is in the input.

This captures my thoughts very well. That's why these images get old very quickly - you can basically imagine the same thing in your mind. There's no real whimsy, surprise or creativity there. For now at least.

Post reply on HN