Live data from Hacker News

How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

datature.io

11–20 of 70 posts

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#11
post #5

With the luggage example it seems to only generate backgrounds where the lighting makes sense? That's kind of interesting. I was wondering how it would handle the highlight on the right.

Giving Stable Diffusion constraints forces it to get creative.

It’s the best argument against “AI generated images are just collages”.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#12

ControlNet model specifically the scribble ControlNet (and ComfyUI) was major gamechanger for me. Was getting good results with just SD and occassional masking but it would take hours and hours to hone in and composite a complex scene with specific requirements & shapes (with most of the work spent currating the best outputs and then blending them into a scene with Gimp/Inkscape). Masking is unintuitive compared to t…

That's for sure - I think we have seen other kind of edge detector or filter work better for differing use cases, especially around foreground images where you want to retain more information (i.e. images with small nitty-gritty details)

In this post, we just seek to showcase the fastest way to do it - and how augmentation may potentially help vary the position!

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#13
post #6
post #3

SD outputs have an "uncanny valley" type of quality to them. You just KNOW when an image is from SD. And I have looked at getting started with SD, but the requirements and setup and +-prompting "language" just kind of turned me off the whole thing. Whereas with DALL E you can get some hyper-realistic images from it with very little effort using plain human language. I guess my point is to ask whether SD is worth both…

One major benefit and the reason why I use the StableDiffusion tools and models is because I can run them at home on my relatively old NVIDIA 2080 GPU with 8GB of VRAM. Costs me nothing (besides electricity). Depends if you value this kind of freedom in life. You can do some things such as colorizing black and white images with the Recolor model. https://huggingface.co/stabilityai/control-lora

Very interesting - thank you for sharing this. Would love to explore this as a team and perhaps put out a blog on helping others get started with control-lora

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#14
post #3

SD outputs have an "uncanny valley" type of quality to them. You just KNOW when an image is from SD. And I have looked at getting started with SD, but the requirements and setup and +-prompting "language" just kind of turned me off the whole thing. Whereas with DALL E you can get some hyper-realistic images from it with very little effort using plain human language. I guess my point is to ask whether SD is worth both…

Try: https://github.com/lllyasviel/Fooocus

I also recommend a good photorealistic base model, like RealVis XL.

In my experience its like DALL E but straight up better, more customizable, and local. And thats before you start trying finetunes and LORAs.

Other UIs will do SDXL, but every one I tried is terrible without all those default fooocus augmentations.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#15
post #5

With the luggage example it seems to only generate backgrounds where the lighting makes sense? That's kind of interesting. I was wondering how it would handle the highlight on the right.

Giving Stable Diffusion constraints forces it to get creative. It’s the best argument against “AI generated images are just collages”.

This is a general result. For example, ChatGPT struggles hard with following lexical, syntactic, or phonetic constraints in prompts due to the tokenization scheme - see https://paperswithcode.com/paper/most-language-models-can-be...

LLMs + Diffusors are super charged when using techniques like constraints, controlnet, regional prompting, and related techniques.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#18
post #3

SD outputs have an "uncanny valley" type of quality to them. You just KNOW when an image is from SD. And I have looked at getting started with SD, but the requirements and setup and +-prompting "language" just kind of turned me off the whole thing. Whereas with DALL E you can get some hyper-realistic images from it with very little effort using plain human language. I guess my point is to ask whether SD is worth both…

Try: https://github.com/lllyasviel/Fooocus I also recommend a good photorealistic base model, like RealVis XL. In my experience its like DALL E but straight up better, more customizable, and local. And thats before you start trying finetunes and LORAs. Other UIs will do SDXL, but every one I tried is terrible without all those default fooocus augmentations.

SDXL is great but it's in no way better than DALL E as far as straight text-to-image goes apart from the lack of censorship.

It has plenty of other advantages, but you can't tell it "make me a cute illustration of a 2 year old girl with Blaze from Blaze and the Monster Machines on a birthday cake with a large 2 candle on it."

DALL E will nail that, more or less. SDXL very much won't.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#19
post #9
post #6

Earlier quoted context omitted.

One major benefit and the reason why I use the StableDiffusion tools and models is because I can run them at home on my relatively old NVIDIA 2080 GPU with 8GB of VRAM. Costs me nothing (besides electricity). Depends if you value this kind of freedom in life. You can do some things such as colorizing black and white images with the Recolor model. https://huggingface.co/stabilityai/control-lora

I mean, I'm running DALLE 3 on a browser from an old laptop and I've generated probably over 15k images in 2 weeks, spanning the gamut from memes to art to lewds (with jailbreaks). The ability to completely scrap what you're building and start totally fresh at the drop of a hat with a new line of ideas and get instant results seems pretty freeing to me.

That’s fine, but it’s like asking: “Why would anyone want to have a personal website when you can just write stuff on Facebook and Twitter and it’s so much easier?”

Stable Diffusion is an open model that you can run locally on your own computer without anyone’s permission. Dall-E is a closed model that runs on OpenAI’s very expensive server farm, and they can change how it works and what it costs whenever they please.

Right now AI is in the Uber-style expansion phase where the service is practically given away to conquer market share. Once the hypergrowth is over, OpenAI will start raising their prices just like Uber did.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#20
post #3

SD outputs have an "uncanny valley" type of quality to them. You just KNOW when an image is from SD. And I have looked at getting started with SD, but the requirements and setup and +-prompting "language" just kind of turned me off the whole thing. Whereas with DALL E you can get some hyper-realistic images from it with very little effort using plain human language. I guess my point is to ask whether SD is worth both…

That's funny to hear because DALL-E 3 mainly improves prompt understanding, it hallucinates like mad with faces and hands, and doesn't seem to do anything to improve them like Midjourney for example.

>Whereas with DALL E you can get some hyper-realistic images from it with very little effort using plain human language.

Hyper-realistic, but is it what you want from it? Are you able to guide it into doing exactly what you want? If you have such requirements that just a natural language prompt is enough and is somehow faster than sketching and providing references, of course use it. I'm not so lucky, I don't get what I want from it, and no amount of prompt understanding will make it easier. Although SD/SDXL doesn't pass the quality bar either, not because it's not "detailed" or "hyper-realistic" enough, but because it doesn't pay attention to the things that should be prioritized, like linework or lighting. Neither does any other model. Controlnets and LoRAs alone aren't sufficient for controllability either, mostly because it's too small to understand high-level concepts. So I don't use anything.

Post reply on HN