Live data from Hacker News

How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

datature.io

21–30 of 70 posts

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#21
post #3

SD outputs have an "uncanny valley" type of quality to them. You just KNOW when an image is from SD. And I have looked at getting started with SD, but the requirements and setup and +-prompting "language" just kind of turned me off the whole thing. Whereas with DALL E you can get some hyper-realistic images from it with very little effort using plain human language. I guess my point is to ask whether SD is worth both…

You just KNOW when an image is from SD

No, you know when a beginner generated an image in Stable Diffusion. With enough skill and attention, you will not.

Sure, there is a learning curve and it takes more time to get to a good result. But in turn, it gives you control far beyond what the competition can offer.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#22
post #3

SD outputs have an "uncanny valley" type of quality to them. You just KNOW when an image is from SD. And I have looked at getting started with SD, but the requirements and setup and +-prompting "language" just kind of turned me off the whole thing. Whereas with DALL E you can get some hyper-realistic images from it with very little effort using plain human language. I guess my point is to ask whether SD is worth both…

> You just KNOW when an image is from SD.

You don't. People think they do, but they don't.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#24
post #3

SD outputs have an "uncanny valley" type of quality to them. You just KNOW when an image is from SD. And I have looked at getting started with SD, but the requirements and setup and +-prompting "language" just kind of turned me off the whole thing. Whereas with DALL E you can get some hyper-realistic images from it with very little effort using plain human language. I guess my point is to ask whether SD is worth both…

I’m assuming you haven’t used SDXL?

Give it a go with invokeAI - you can create images that I guarantee you wouldn’t know were generated. Like anything (photography included) it’s a skill.

Examples:

  - https://civitai.com/images/2862100
  - https://civitai.com/images/2339666
  - https://civitai.com/images/2846876

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#25
post #9
post #6

Earlier quoted context omitted.

One major benefit and the reason why I use the StableDiffusion tools and models is because I can run them at home on my relatively old NVIDIA 2080 GPU with 8GB of VRAM. Costs me nothing (besides electricity). Depends if you value this kind of freedom in life. You can do some things such as colorizing black and white images with the Recolor model. https://huggingface.co/stabilityai/control-lora

I mean, I'm running DALLE 3 on a browser from an old laptop and I've generated probably over 15k images in 2 weeks, spanning the gamut from memes to art to lewds (with jailbreaks). The ability to completely scrap what you're building and start totally fresh at the drop of a hat with a new line of ideas and get instant results seems pretty freeing to me.

With SD I can generate at least 15k images daily on my old laptop, I can train it with new styles, characters, real people, etc.; download thousands of new styles, characters, real people, etc. from Civitai, and best of all, never worry about ever losing access to it, being censored, having to jailbreak it, being snooped on, etc.

Plus a million other tools that the community has made for it, like ControlNet or things like AnimateDiff to create videos. I can also easily create all kinds of scripts and workflows.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#26
post #24
post #3

SD outputs have an "uncanny valley" type of quality to them. You just KNOW when an image is from SD. And I have looked at getting started with SD, but the requirements and setup and +-prompting "language" just kind of turned me off the whole thing. Whereas with DALL E you can get some hyper-realistic images from it with very little effort using plain human language. I guess my point is to ask whether SD is worth both…

I’m assuming you haven’t used SDXL? Give it a go with invokeAI - you can create images that I guarantee you wouldn’t know were generated. Like anything (photography included) it’s a skill. Examples: - https://civitai.com/images/2862100 - https://civitai.com/images/2339666 - https://civitai.com/images/2846876

At glance I get uncanny valley from two. After looking closer it's likely because with the photo of the couple at the cinema, the woman's arm around him is wearing the wrong clothes. Then photo of the guy with a hat, his neck piece is asymmetric.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#27
post #5

With the luggage example it seems to only generate backgrounds where the lighting makes sense? That's kind of interesting. I was wondering how it would handle the highlight on the right.

In ComfyUI you could run the image through a style-to-style (sdxl refinement might even pull it off) model to change the lighting without changing the content. Or use another ControlNet. Your workflow can get arbitrarily complex.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#28
post #3

SD outputs have an "uncanny valley" type of quality to them. You just KNOW when an image is from SD. And I have looked at getting started with SD, but the requirements and setup and +-prompting "language" just kind of turned me off the whole thing. Whereas with DALL E you can get some hyper-realistic images from it with very little effort using plain human language. I guess my point is to ask whether SD is worth both…

You know when they're bad enough that you know.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#29
post #24

Earlier quoted context omitted.

I’m assuming you haven’t used SDXL? Give it a go with invokeAI - you can create images that I guarantee you wouldn’t know were generated. Like anything (photography included) it’s a skill. Examples: - https://civitai.com/images/2862100 - https://civitai.com/images/2339666 - https://civitai.com/images/2846876

At glance I get uncanny valley from two. After looking closer it's likely because with the photo of the couple at the cinema, the woman's arm around him is wearing the wrong clothes. Then photo of the guy with a hat, his neck piece is asymmetric.

The eyes are messed up in the second one too which instantly gives it away

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#30
post #7
post #3

SD outputs have an "uncanny valley" type of quality to them. You just KNOW when an image is from SD. And I have looked at getting started with SD, but the requirements and setup and +-prompting "language" just kind of turned me off the whole thing. Whereas with DALL E you can get some hyper-realistic images from it with very little effort using plain human language. I guess my point is to ask whether SD is worth both…

DALL-E within ChatGPT uses GPT-4 to rewrite what you ask for into a good text-to-image prompt. You could probably do something similar with Stable Diffusion with just a little upfront effort tuning that system prompt.

Somewhat, but dalle3 is hugely better at understanding a description and relationships.
Post reply on HN