Live data from Hacker News

Try Stable Diffusion's Img2Img Mode

huggingface.co

71–80 of 166 posts

Re: Try Stable Diffusion's Img2Img Mode

#71
post #13

If you have a GPU with >4GB of VRAM and you want to run this locally, here's a fork of the Stable Diffusion repo with a convenient web UI: https://github.com/hlky/stable-diffusion It supports both txt2img and img2img. (Not affiliated.) Edit: Incidentally, I tried running it on a CPU. It is possible, but it took 3 minutes instead of 10 seconds to produce an image. It also required me to hack up the script in a really…

What kind of GPU do you have? It takes several minutes to produce an image on my 1070.

Takes 3 minutes (for a prompt resulting in a set of 4 images) on my 1080 as well. Really astonished that it takes GP about the same time using just a CPU. Seems like the older generation of GPUs isn't much better than CPUs in regards to ML stuff.

Re: Try Stable Diffusion's Img2Img Mode

#72
post #13

If you have a GPU with >4GB of VRAM and you want to run this locally, here's a fork of the Stable Diffusion repo with a convenient web UI: https://github.com/hlky/stable-diffusion It supports both txt2img and img2img. (Not affiliated.) Edit: Incidentally, I tried running it on a CPU. It is possible, but it took 3 minutes instead of 10 seconds to produce an image. It also required me to hack up the script in a really…

What kind of GPU do you have? It takes several minutes to produce an image on my 1070.

To add another data point, my GTX1080 takes ~60 sec to generate a pair of 500x500 images using txt2img. Haven't tried img2img yet as the UI package I went with is a bit buggy with it

Re: Try Stable Diffusion's Img2Img Mode

#73
post #67

I've been trying to get some sensible images out of my descriptions, but I fail miserably. In this case I had the prompt "cow chewing bone" with 4 squares representing the two pair of feet, the body and the head. None cared about chewing on a bone. With DALL·E 2 I tried to get an image of a little girl building sandcastles and a monster threatening her: "little scared girl building a sandcastle and a big angry monste…

I managed to get one that was correct with "A little scared girl is building a sandcastle, while a monster is looking at her. Award-winning photograph.", but I couldn't figure out a phrasing where it wouldn't most of the time get confused thinking that the sandcastle is the monster, or that the girl is the monster.

Re: Try Stable Diffusion's Img2Img Mode

#74
post #13

If you have a GPU with >4GB of VRAM and you want to run this locally, here's a fork of the Stable Diffusion repo with a convenient web UI: https://github.com/hlky/stable-diffusion It supports both txt2img and img2img. (Not affiliated.) Edit: Incidentally, I tried running it on a CPU. It is possible, but it took 3 minutes instead of 10 seconds to produce an image. It also required me to hack up the script in a really…

What kind of GPU do you have? It takes several minutes to produce an image on my 1070.

also on a 1070, I can generate an image in ~15 seconds, surely you're doing something wrong.

Re: Try Stable Diffusion's Img2Img Mode

#75
post #48

Can't help but see this as a friendly herald for a dystopian future.

There you go. We were promised quite far fetched things like 'flying cars', 'time travel', 'life extension' and 'universal basic income'. Instead we get Deep Learning AI's being used for generating and faking sentences, images, videos, voices, code and digital art all being trained on mountains of data in data centers all significantly contributing to the already burning up of the planet to no benefit and no efficien…

Tune out the news and look for a futurism themed news source[0], and you can get your techno-optimistic pipe dreams back, right now. What you describe as "we were promised" and "we got" is just you seeing a bit more of the world. It was filled with empty promises and horrors in the beforetimes too, today's flavor is nothing special - just in its specifics for today's culture. And there's also plenty of people doing great things in the past too, perhaps not promising those things, just living their life doing something they believe in, and getting results.

My point being, there's enough good things and bad things to fit any mood and worldview. Everyone can basically pick as they'd like.

[0] something like this https://www.reddit.com/r/Futurology/

Re: Try Stable Diffusion's Img2Img Mode

#76
I've been playing with this for a few hours. It's slow going -- you really need a fast GPU with a lot of RAM to make this very usable.

I ended up paying the $10 for Google Colab Pro and that's how I've been using this. Maybe I'll figure out how to get this working on my old 1080 TI to see if it's faster.

Anyway, for the one that I'm using which has a web UI, you can use this Colab link. It's pretty great! https://colab.research.google.com/drive/1KeNq05lji7p-WDS2BL-...

What I really wish was that the img2img tool could be used to take a text2img output and then "refine" it further. As it is, the img2img tool doesn't seem particularly great.

People on Reddit are talking about "I just generate 100 images and pick the best one"... but this is incredibly slow on the P100 GPU that Google has me on. Does this just require a monster GPU like a 3080/3090 in order to get any decent results?

Re: Try Stable Diffusion's Img2Img Mode

#77
post #70

Earlier quoted context omitted.

You must be kidding it’s hard labor creating images and you are complaining that it’s going to be easier in the future? You are complaining that the skill gap is eradicated and it now only depends on your ideas to produce compelling work? That’s absolutely insane everything that takes labor intensive tasks and makes them basically free is the future. This is not dystopian it’s our future, a future in which you and yo…

The skill gap was what made it valuable. There is not that much value in flooding everyone with generated images that have no story, no person behind it, no effort, no reason to exist, out of place. “Art” is as much about the final image as it is about the person who made it, why she made it and how she got there. Also there is pretty dystopian angle because these tools balantly stole all the work of artists by “lear…

You in 1894: Why does no one think about all the horses and stable boys. The value of art is also not in the hours you put into creating it it's about the idea behind it and how it's conveyed. Artists are also always stealing especially concept artists the only thing to say is "photobashing". You talk like someone who has absolutely no idea how the sausage is made.

Re: Try Stable Diffusion's Img2Img Mode

#79
post #74

Earlier quoted context omitted.

What kind of GPU do you have? It takes several minutes to produce an image on my 1070.

also on a 1070, I can generate an image in ~15 seconds, surely you're doing something wrong.

What settings are you using it the people with 1080s above you are taking 3 minutes?
Post reply on HN