Live data from Hacker News

How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

datature.io

41–50 of 70 posts

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#41

ControlNet model specifically the scribble ControlNet (and ComfyUI) was major gamechanger for me. Was getting good results with just SD and occassional masking but it would take hours and hours to hone in and composite a complex scene with specific requirements & shapes (with most of the work spent currating the best outputs and then blending them into a scene with Gimp/Inkscape). Masking is unintuitive compared to t…

any tutorial you would recommend? I found https://comfyanonymous.github.io/ComfyUI_examples/controlnet...

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#42
post #24

Earlier quoted context omitted.

I’m assuming you haven’t used SDXL? Give it a go with invokeAI - you can create images that I guarantee you wouldn’t know were generated. Like anything (photography included) it’s a skill. Examples: - https://civitai.com/images/2862100 - https://civitai.com/images/2339666 - https://civitai.com/images/2846876

At glance I get uncanny valley from two. After looking closer it's likely because with the photo of the couple at the cinema, the woman's arm around him is wearing the wrong clothes. Then photo of the guy with a hat, his neck piece is asymmetric.

That's still a 1 in 3 success rate, at the cost of writing a prompt and waiting a minute.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#43

ControlNet model specifically the scribble ControlNet (and ComfyUI) was major gamechanger for me. Was getting good results with just SD and occassional masking but it would take hours and hours to hone in and composite a complex scene with specific requirements & shapes (with most of the work spent currating the best outputs and then blending them into a scene with Gimp/Inkscape). Masking is unintuitive compared to t…

You don't need to go through Gimp or Inkscape, this is built-in to the auto1111 ControlNet UI. You just dump the existing photo there and you can select a bunch of pre-processors like edge-detection or 3D depth extraction, which is then fed into ControlNet to generate a new image.

This is super powerful for example visualizing the renovation of an apartment room or house exterior.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#45
post #36

The versatility of Stable Diffusion, especially when combined with tools like ControlNet, highlights the advantages of a more controlled image generation process. While DALL-E and others provide ease and speed, the depth of customization and local processing capabilities of SD models cater to those seeking deeper creative control and independence.

It is interesting isn't it? Because we have "AI" generating the image, but we still seem to want to "paint" or have control over the creative process. Prompts seem to be a new type of camera, lens or paintbrush.

There's at least three "levels" you can consider with image generation: composition, facial likeness and style. Prompts are pretty weak at composition and are the strongest point of controlnets - they do a great deal to make up for the weakness. But there are some compositions SD can't find even when given detailed controlnets.

Style generality is frequently lost in fine-tuned models. The original dreambooth tried to get around this by generating lots of images of the class to retain generality, but it's time intensive to generate all the extra images (and ideally do some QC on them) and train on them too, so it's not often done.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#46
Love all this AI stuff, would love to play more with it, but sadly I'm on a 2015 iMac, great for everything else I do but can't do this stuff.

It's pricey to get a windows machine + GPU and the cloud options seem a bit more limited and add up quickly too, but it is amazing tech.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#47

Love all this AI stuff, would love to play more with it, but sadly I'm on a 2015 iMac, great for everything else I do but can't do this stuff. It's pricey to get a windows machine + GPU and the cloud options seem a bit more limited and add up quickly too, but it is amazing tech.

I have done a bunch of stable diffusion stuff on colab. The free version works if you are lucky enough to get a GPU. Used to happen more often before. But the premium colab isn't badly priced either.

Here is a colab link to open comfyUI

https://github.com/FurkanGozukara/Stable-Diffusion/blob/main...

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#48

If I have a large dataset or photos with my face? Can I generate my own images in different places and environments using this?

Yep. Lora's are the easiest way to go. Loads of tutorials on Youtube. This is a good one: https://www.youtube.com/watch?v=70H03cv57-o

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#49
post #48

If I have a large dataset or photos with my face? Can I generate my own images in different places and environments using this?

Yep. Lora's are the easiest way to go. Loads of tutorials on Youtube. This is a good one: https://www.youtube.com/watch?v=70H03cv57-o

Looks like it requires good hardware to run this? My GPU is too old for this.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#50
post #18

Earlier quoted context omitted.

Try: https://github.com/lllyasviel/Fooocus I also recommend a good photorealistic base model, like RealVis XL. In my experience its like DALL E but straight up better, more customizable, and local. And thats before you start trying finetunes and LORAs. Other UIs will do SDXL, but every one I tried is terrible without all those default fooocus augmentations.

SDXL is great but it's in no way better than DALL E as far as straight text-to-image goes apart from the lack of censorship. It has plenty of other advantages, but you can't tell it "make me a cute illustration of a 2 year old girl with Blaze from Blaze and the Monster Machines on a birthday cake with a large 2 candle on it." DALL E will nail that, more or less. SDXL very much won't.

Heh, yeah, that is true

https://ibb.co/1m0bLWC

More cherrypicking and messing with styles is getting closer, but nothing like Dall-E's first try I'm sure.

Post reply on HN