Live data from Hacker News

How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

datature.io

31–40 of 70 posts

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#31
post #24
post #3

SD outputs have an "uncanny valley" type of quality to them. You just KNOW when an image is from SD. And I have looked at getting started with SD, but the requirements and setup and +-prompting "language" just kind of turned me off the whole thing. Whereas with DALL E you can get some hyper-realistic images from it with very little effort using plain human language. I guess my point is to ask whether SD is worth both…

I’m assuming you haven’t used SDXL? Give it a go with invokeAI - you can create images that I guarantee you wouldn’t know were generated. Like anything (photography included) it’s a skill. Examples: - https://civitai.com/images/2862100 - https://civitai.com/images/2339666 - https://civitai.com/images/2846876

I can see at least three finger issues with the couple in the cinema.

More than that though: I use SDXL quite a bit for fun, and while I like it and it can be very good, it's still prone to getting stuck in a David Cronenberg mode for reasons I can't solve.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#32
The versatility of Stable Diffusion, especially when combined with tools like ControlNet, highlights the advantages of a more controlled image generation process. While DALL-E and others provide ease and speed, the depth of customization and local processing capabilities of SD models cater to those seeking deeper creative control and independence.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#34
post #3

SD outputs have an "uncanny valley" type of quality to them. You just KNOW when an image is from SD. And I have looked at getting started with SD, but the requirements and setup and +-prompting "language" just kind of turned me off the whole thing. Whereas with DALL E you can get some hyper-realistic images from it with very little effort using plain human language. I guess my point is to ask whether SD is worth both…

Dall-E has the same problems as other models. Try generating a clockwork mechanism with it, for example.

SD is worth bothering with because it's open, you can run and extend it yourself.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#35
post #31
post #24

Earlier quoted context omitted.

I’m assuming you haven’t used SDXL? Give it a go with invokeAI - you can create images that I guarantee you wouldn’t know were generated. Like anything (photography included) it’s a skill. Examples: - https://civitai.com/images/2862100 - https://civitai.com/images/2339666 - https://civitai.com/images/2846876

I can see at least three finger issues with the couple in the cinema. More than that though: I use SDXL quite a bit for fun, and while I like it and it can be very good, it's still prone to getting stuck in a David Cronenberg mode for reasons I can't solve.

Oh yeah it’s not one-shot perfect but it gets you 90% of the way there for a lot of things. I’m super impressed with it.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#36

The versatility of Stable Diffusion, especially when combined with tools like ControlNet, highlights the advantages of a more controlled image generation process. While DALL-E and others provide ease and speed, the depth of customization and local processing capabilities of SD models cater to those seeking deeper creative control and independence.

It is interesting isn't it? Because we have "AI" generating the image, but we still seem to want to "paint" or have control over the creative process.

Prompts seem to be a new type of camera, lens or paintbrush.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#37
post #18

Earlier quoted context omitted.

Try: https://github.com/lllyasviel/Fooocus I also recommend a good photorealistic base model, like RealVis XL. In my experience its like DALL E but straight up better, more customizable, and local. And thats before you start trying finetunes and LORAs. Other UIs will do SDXL, but every one I tried is terrible without all those default fooocus augmentations.

SDXL is great but it's in no way better than DALL E as far as straight text-to-image goes apart from the lack of censorship. It has plenty of other advantages, but you can't tell it "make me a cute illustration of a 2 year old girl with Blaze from Blaze and the Monster Machines on a birthday cake with a large 2 candle on it." DALL E will nail that, more or less. SDXL very much won't.

SD XL understands prompts much better than 1.5. So the next version of SD might be comparable to Dall-E without censorship.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#38

ControlNet model specifically the scribble ControlNet (and ComfyUI) was major gamechanger for me. Was getting good results with just SD and occassional masking but it would take hours and hours to hone in and composite a complex scene with specific requirements & shapes (with most of the work spent currating the best outputs and then blending them into a scene with Gimp/Inkscape). Masking is unintuitive compared to t…

Comfy ui is really nice. The fact that the node graph is saved as png metadata actually makes node based workflows super fluent and reproducible since all you need to do to get the graph for your image is to drag and drop the result png to the gui. This feels like a huge quality-of-life improvement compared to any other lightweight node tools I’ve used.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#39
post #3

SD outputs have an "uncanny valley" type of quality to them. You just KNOW when an image is from SD. And I have looked at getting started with SD, but the requirements and setup and +-prompting "language" just kind of turned me off the whole thing. Whereas with DALL E you can get some hyper-realistic images from it with very little effort using plain human language. I guess my point is to ask whether SD is worth both…

Depends what you want.

Dalle 3 is super good, but lacks the creative control controlnets and ip-adapter provide. So for instance afaik there is no way to perform style transfers, or ’paint a van gogh portrait over my pencil sketch’.

Both are good currently but at different things.

”Prompt engineering” is and will be total bs. Dalle3/chatgpt provides the actual workflow we want where we describe to the intelligent agent (chatpgt) what we want and it worries over the accidental-complexity-intricasies of the clip model itself.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#40
post #3

SD outputs have an "uncanny valley" type of quality to them. You just KNOW when an image is from SD. And I have looked at getting started with SD, but the requirements and setup and +-prompting "language" just kind of turned me off the whole thing. Whereas with DALL E you can get some hyper-realistic images from it with very little effort using plain human language. I guess my point is to ask whether SD is worth both…

I can run Stable Diffusion on my local machine. It is open source and weights are public, giving me in theory access to anything I want to modify.

I cant change anything on DALL E, I can just take the input or change the prompt.

Also it is a centralized service that can be shut down, modified, censored or become very expensive at any time.

Post reply on HN