Live data from Hacker News

Stable Diffusion is a big deal

simonwillison.net

471–480 of 488 posts

Re: Stable Diffusion is a big deal

#471
post #175

> Stable Diffusion has been trained on millions of copyrighted images scraped from the web. My brain has been trained on even more copyrighted material. Every book I read, every tv show I watch, the toys I played with as a child. It's hard to imagine that I could come up with anything that is not inspired by copyrighted work.

> My brain has been trained on even more copyrighted material

What your brain has learned cannot be transferred with an USB stick in seconds

Not even your offspring will receive any of it

If I want to learn everything you know, I have to learn what you learned , assuming I will be able to

Kinda of a big difference, don't you think?

These kinds of comments are embarrassingly low effort, just because we threw rocks at each other when we all were chimps, doesn't mean that guns haven't been a game changer and made homicide easier even for people that would never be able to hit anybody by throwing rocks.

Re: Stable Diffusion is a big deal

#472
post #452
post #246

I've said this ad nauseam - but people who think this is going to kill an industry clearly have no idea of said industry. It's a fantastically great tool, and a very exciting space, but reducing the function of creatives to people who draw pretty pictures is staggeringly ignorant.

You pose it like an all or nothing equation, which it isn't. The way I see it, if you'd consider the art world a pyramid, the bottom is about to fall out. A lot of commercial artwork serves no deeper meaning but pretty decoration. The emphasis will move to ideas instead of just execution. Artists will soon find out about the avalanche of people that have creative ideas yet can't draw or paint. They'll be unlocked.

In context it is all or nothing - because people on HN think illustrators and artists are what they find on fiverr.com. They're making hugely naive blanket statements saying that this will destroy an industry and make creatives unemployable. These doomsayers have literally not the faintest idea about the job they think is being erased by txt2img.

If people were saying "oh hey this is going to give the lazystock on iStockPhoto a run for their money", I wouldn't debate that point, it's true. However that's not the industry, and it's certainly not where the money is - neither in # of customers nor total spend. Those people who you might think are customers simply put: aren't, they get by with images stolen from google images and bundled clipart, or frankly: nothing at all.

Now this isn't to say that txt2img isn't useful or exciting. I can say that it is the largest and most significant expansion of creative tech since the advent of DTP. This will absolutely accelerate and open the door to not just higher standards, but new ways of rapidly ideating concepts. I've already seen fantastic examples of txt-to-image-to-mesh-to-live animation. All automated through AI.

This is also why I speak against the other kinds of naysayers: the ones that think this tech is unimportant. These types are being incredibly short sighted and acting like we're looking at this tech's endpoint, rather than its infancy.

tl,dr: No creatives are not being put out of the job. Yes this tech is incredibly important.

Re: Stable Diffusion is a big deal

#473

Earlier quoted context omitted.

>You also need some overlap between what the model has likely been trained on and the expected output. Ooooo. Ouch. That's... Kind of a death knell for a good tool in my experience, and suggests the tool in question is a glorified search engine in a sense. With a really confusing query syntax composed of words, and graphical starting states. If I have to become to get anything done... Why not just learn to draw/hire…

Learning to paint hyper realistic paintings is something that takes years if not decades of hard work. Learning how to formulate your queries in a way that the algorithm outputs what you want takes days at worst. If you want something unique and abstract, you're going to need to go through a lot of trial and error to get what you want. That's still a lot easier than teaching yourself how to create such art. "Graphica…

Took a look.

Trollface: Disambiguation of lines in the source image is poorly executed. The model appears confused as to whether those lines are indicative of depth, or lighting artifacts. The shape and perspective are poorly chosen, and in all the resulting images the lighting arrangement is quite inconsistent.

The ears are completely unspecified, so too the nose. This is somewhat of a deliberate omission in a trollface, and adding them in without careful thought as to how it changes the piece is... Well, not the best move. The eyes are terribly arranged in all submissions.

The plate of meat, fries, and beans. You can barely see the beans in the first sample, they are hidden underneath the fries, enough that an inattentive eye may miss them entirely. No specifi ation was given as to the state of the meat, or kind, so I suppose the being cut is a nice bonus. Interesting in a sense since one may get the impression the model may have confused the grammatical deep structure such that "fries and beans" was taken as a compound predicate.

The second with the meat surrounded by the beans is an interesting contrast, but without more samples, I have questions about why all the curated samples include rare beef instead of say, sausage.

The Colloseum: I too could use Photoshop, and select a particular palate. The more interesting aspect here seems to be the color pallete processing, and I'll admit that I wasn't able to find source works of the artist being initated to compare against. Still looking for those.

The Unicorn/Butterfly: These still disturb me in the sense that once again, we're replacing actual artistic technique, with the ability to tweak prompts or assemble graphical starting states/prompt combos. Is it making some hellish form of combined Natural Language/graphical programming pipeline? Yes.

However, none of this would have any value without being trained on works done by previous artists who likely were not asked whether or not they wanted their works included in the dataset.

As the guy who blew up a Philosophy of Art class by positing that a well executed forgery was as much a work of Art as the imitated piece, I still see here more problems than solutions. Yes, a new art form may have emerged. However, with it comes serious questions around data curation practices. As for the efficacy of the model/runtime characteristics/how this bodes for the environment... I'm increasingly concerned the more I apply ny "what if everyone started doing this?" supposition.

In short, see a hell of a lot of hype, but precious lite coming to terms with what will ultimately be the hard questions.

Re: Stable Diffusion is a big deal

#474

This doesn't strictly have to do with stable diffusion, but I really really hope that someone eventually creates a service that hosts live interfaces for all of the legacy image generation models. The turn around on these things is wild, and I worry that unique image generation techniques will become antiquated and forgotten.

In case with SD, you can download every version yourself. The main code is in git (you can go back to any point), and weights are versioned, I have every version saved, for example.

Re: Stable Diffusion is a big deal

#475
post #472
post #452

Earlier quoted context omitted.

You pose it like an all or nothing equation, which it isn't. The way I see it, if you'd consider the art world a pyramid, the bottom is about to fall out. A lot of commercial artwork serves no deeper meaning but pretty decoration. The emphasis will move to ideas instead of just execution. Artists will soon find out about the avalanche of people that have creative ideas yet can't draw or paint. They'll be unlocked.

In context it is all or nothing - because people on HN think illustrators and artists are what they find on fiverr.com. They're making hugely naive blanket statements saying that this will destroy an industry and make creatives unemployable. These doomsayers have literally not the faintest idea about the job they think is being erased by txt2img. If people were saying "oh hey this is going to give the lazystock on iS…

I agree with you. I think it's understandable that non-artists commonly associate art with what they interact with or see the most: illustration and decoration, stuff found at artstation, the like.

Surely you have a point that this does not cover the entire world of art, but I think it would be helpful if you constructively explain which parts are less or not affected, instead of calling people ignorant.

Re: Stable Diffusion is a big deal

#476
post #455

Earlier quoted context omitted.

Not sure what you mean by "standards/frameworks". It's an art rather than a science. Some of the tips passed around smack somewhat of cargo-culting, some genuinely make a difference. Everyone has their own approach and some techniques work for one topic but not others.

There's even a brand new word for it: prompt superstition. "8k, highly detailed, ultra super very detailed, epic lighting, octane render"

Just came across a nice post on the topic of prompt superstition: https://old.reddit.com/r/StableDiffusion/comments/x32c2o/sta...

Re: Stable Diffusion is a big deal

#477
post #455

Earlier quoted context omitted.

There's even a brand new word for it: prompt superstition. "8k, highly detailed, ultra super very detailed, epic lighting, octane render"

Just came across a nice post on the topic of prompt superstition: https://old.reddit.com/r/StableDiffusion/comments/x32c2o/sta...

Very interesting, thanks for sharing.

Re: Stable Diffusion is a big deal

#478

Earlier quoted context omitted.

Without trained models from human creativity, what can AI do ? These AI emerged because of human creativity. Picasso created it’s new art form from its own creativity. He created something no one ever though of before. Now AI are fueled with Picasso’s drawing and can produce art that looks like his maybe. But what about creating something entirely new that has never been fueled into the engine. Could the AI invent so…

I think it's undeniable that AI can create novel things, the question is if AI can create novel things that are also interesting. A randomized 600x600 png is novel, but it isn't at all interesting, much of what goes into making a piece of art interesting is not a quantifiable or well-defined goal. That's not to say that AI is better than humans, just the opposite, art is a deeply human object, and I do not know if we…

I see your point. We agree that on a 600x600png there is a finite number of possibilities given a set number of colors. It could be possible to « brute force » all the images possible which would be the equivalent of trying all the combinations of letters to write a book. So creating things is not really a problem. Creating relevant things from a purpose is.

The fact is, we know no other intelligence except ours. We are the only known being that appreciate art or books, from our subjective conscience. Can AI create its own art, understood and appreciated by itself ? Has intelligence a meaning without « humans » ? It’s part of the same thing. So the AI is modeled after ours. So it’s not its own thing and cannot understand what it creates and if what is created is relevant. AI is just an amazingly powerful tool.

Re: Stable Diffusion is a big deal

#479
post #475
post #472

Earlier quoted context omitted.

In context it is all or nothing - because people on HN think illustrators and artists are what they find on fiverr.com. They're making hugely naive blanket statements saying that this will destroy an industry and make creatives unemployable. These doomsayers have literally not the faintest idea about the job they think is being erased by txt2img. If people were saying "oh hey this is going to give the lazystock on iS…

I agree with you. I think it's understandable that non-artists commonly associate art with what they interact with or see the most: illustration and decoration, stuff found at artstation, the like. Surely you have a point that this does not cover the entire world of art, but I think it would be helpful if you constructively explain which parts are less or not affected, instead of calling people ignorant.

I’ll definitely continue to call people ignorant when they make grand unsubstantiated claims, that frankly are nothing more than trollish internet behaviour.

A better approach for people is to ask questions, rather than trying to write controversial falsehoods.

Right now there is at least one high ranking submissions on HN where a creative details how this won’t end their career, but you don’t need to read it - social media is filled with creatives literally rejoicing - no one is sweating this.

So to that: I say that ignorance to this is definitely a choice.

Re: Stable Diffusion is a big deal

#480

Earlier quoted context omitted.

A few months ago this task was virtually impossible. Then it was possible, but extremely expensive and pay-walled behind the "Open"AI' website. As of about a week ago this tech runs on consumer GPUs. The weights have been downloaded 100s of thousands of times, and fine-tuning / modifying is possible. Training from scratch is about $500k still, but it will only get cheaper and easier.

This doesn't contradict anything I have written. The average technical user will be unable to train this exact model (not to mention the supposed future more powerful ones) in their basement in this decade.

That just feels like such a pessimistic forecast to me. Of course, the current trajectory of improvements in model efficiency and better commercial GPUs / ML-accelerators may hit a wall.

But I would not be surprised if this was trainable on a commercial GPU at home within that time. But I think another important trend that we are seeing is that you don't need to train these models from scratch.

Open-source "foundation models" means that you can usually get away with the much easier task of fine-tuning, as to not throw away / re-learn everything that these large models have already fit.

Edit: I initially said 2-5 years, but on more reflection this does seem optimistic (for training from scratch).

Post reply on HN