Live data from Hacker News

How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

datature.io

51–60 of 70 posts

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#51
post #9
post #6

Earlier quoted context omitted.

One major benefit and the reason why I use the StableDiffusion tools and models is because I can run them at home on my relatively old NVIDIA 2080 GPU with 8GB of VRAM. Costs me nothing (besides electricity). Depends if you value this kind of freedom in life. You can do some things such as colorizing black and white images with the Recolor model. https://huggingface.co/stabilityai/control-lora

I mean, I'm running DALLE 3 on a browser from an old laptop and I've generated probably over 15k images in 2 weeks, spanning the gamut from memes to art to lewds (with jailbreaks). The ability to completely scrap what you're building and start totally fresh at the drop of a hat with a new line of ideas and get instant results seems pretty freeing to me.

I'm using Dall-e 3 through ChatGPT but it seems to limit the amount of images I can generate per half hour. I haven't figured out the actual limit but sometimes I go to generate an image and it just says "You've reached your image generation cap please wait _n_ minutes before trying again"

Are you getting around that somehow? Even if it'll let me generate 36 images per half hour (which seems like it's probably lower than that) I can only generate 6k in 2 weeks prompting 24/7. I'm not scrutinizing your numbers I'm more hoping I'm missing some way to not have to be capped. I already pay for GPT+

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#52
post #48

Earlier quoted context omitted.

Yep. Lora's are the easiest way to go. Loads of tutorials on Youtube. This is a good one: https://www.youtube.com/watch?v=70H03cv57-o

Looks like it requires good hardware to run this? My GPU is too old for this.

You typically want an Nvidia GPU with at least 8GB of VRAM to get started with this stuff. You can get away with less, but it will be slow going. More VRAM is better.

If you're serious about learning how to use these tools, it's far more affordable to rent a GPU in the cloud. Google Colab even has some free tiers with limited access to GPUs that are significantly more powerful than what you would normally put in a desktop.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#53
post #10

I'm building CushyStudio https://github.com/rvion/cushystudio#readme to make Stable Diffusion practical and fun to play with. It's still a bit rough around the corners, and I haven't properly launched it yet, but if you want to play with ControlNets, pre-processors, IP adapters, and all those various SD technologies, it's a pretty fun tool ! I personally use for real-time scribble to image, things like this :) (will…

Looking forward to your launch, I found cushystudio awhile back (maybe from HN?) and cannibalized some of the type generation code to make my own API wrapper for personal uses. Thanks!

I barely got it working in that early alpha but it was super helpful for me as a reference. I'll give it another go now that it's further along, it seemed very promising and I liked your workflow approach

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#54

Love all this AI stuff, would love to play more with it, but sadly I'm on a 2015 iMac, great for everything else I do but can't do this stuff. It's pricey to get a windows machine + GPU and the cloud options seem a bit more limited and add up quickly too, but it is amazing tech.

I have done a bunch of stable diffusion stuff on colab. The free version works if you are lucky enough to get a GPU. Used to happen more often before. But the premium colab isn't badly priced either. Here is a colab link to open comfyUI https://github.com/FurkanGozukara/Stable-Diffusion/blob/main...

They blocked this now on the free version of colab sadly.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#55
post #9

Earlier quoted context omitted.

I mean, I'm running DALLE 3 on a browser from an old laptop and I've generated probably over 15k images in 2 weeks, spanning the gamut from memes to art to lewds (with jailbreaks). The ability to completely scrap what you're building and start totally fresh at the drop of a hat with a new line of ideas and get instant results seems pretty freeing to me.

I'm using Dall-e 3 through ChatGPT but it seems to limit the amount of images I can generate per half hour. I haven't figured out the actual limit but sometimes I go to generate an image and it just says "You've reached your image generation cap please wait _n_ minutes before trying again" Are you getting around that somehow? Even if it'll let me generate 36 images per half hour (which seems like it's probably lower…

When you run out of boost tokens, if you clear your bing search history and restart your browser you get fresh boost tokens. I've been able to do this endlessly. Also if you use it during non-peak hours the wait times are usually 30 seconds or under for 3-4 image generations.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#56
post #55

Earlier quoted context omitted.

I'm using Dall-e 3 through ChatGPT but it seems to limit the amount of images I can generate per half hour. I haven't figured out the actual limit but sometimes I go to generate an image and it just says "You've reached your image generation cap please wait _n_ minutes before trying again" Are you getting around that somehow? Even if it'll let me generate 36 images per half hour (which seems like it's probably lower…

When you run out of boost tokens, if you clear your bing search history and restart your browser you get fresh boost tokens. I've been able to do this endlessly. Also if you use it during non-peak hours the wait times are usually 30 seconds or under for 3-4 image generations.

Sweeeet! Thank you!

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#57

ControlNet model specifically the scribble ControlNet (and ComfyUI) was major gamechanger for me. Was getting good results with just SD and occassional masking but it would take hours and hours to hone in and composite a complex scene with specific requirements & shapes (with most of the work spent currating the best outputs and then blending them into a scene with Gimp/Inkscape). Masking is unintuitive compared to t…

Is there any solution for consistency yet that goes beyond form and structure and gets things like outfits, color, and facial features consistent in an easy way to compose scenes with multiple consistent characters?

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#59

ControlNet model specifically the scribble ControlNet (and ComfyUI) was major gamechanger for me. Was getting good results with just SD and occassional masking but it would take hours and hours to hone in and composite a complex scene with specific requirements & shapes (with most of the work spent currating the best outputs and then blending them into a scene with Gimp/Inkscape). Masking is unintuitive compared to t…

Is there any solution for consistency yet that goes beyond form and structure and gets things like outfits, color, and facial features consistent in an easy way to compose scenes with multiple consistent characters?

LoRA for specific items/characters + regional prompting covers a lot of that area.

Re: How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion

#60
post #7
post #3

SD outputs have an "uncanny valley" type of quality to them. You just KNOW when an image is from SD. And I have looked at getting started with SD, but the requirements and setup and +-prompting "language" just kind of turned me off the whole thing. Whereas with DALL E you can get some hyper-realistic images from it with very little effort using plain human language. I guess my point is to ask whether SD is worth both…

DALL-E within ChatGPT uses GPT-4 to rewrite what you ask for into a good text-to-image prompt. You could probably do something similar with Stable Diffusion with just a little upfront effort tuning that system prompt.

> You could probably do something similar with Stable Diffusion with just a little upfront effort tuning that system prompt.

And, indeed, someone has:

https://github.com/sayakpaul/caption-upsampling

Post reply on HN