Live data from Hacker News

Try Stable Diffusion's Img2Img Mode

huggingface.co

61–70 of 166 posts

Re: Try Stable Diffusion's Img2Img Mode

#61
post #48

Can't help but see this as a friendly herald for a dystopian future.

There you go. We were promised quite far fetched things like 'flying cars', 'time travel', 'life extension' and 'universal basic income'. Instead we get Deep Learning AI's being used for generating and faking sentences, images, videos, voices, code and digital art all being trained on mountains of data in data centers all significantly contributing to the already burning up of the planet to no benefit and no efficien…

People are too pessimistic.

If we had time travel, they’d say now someone can kill our granddads and rewrite history. If we had flying cars, they’d say now our flats are worthless because everyone can fly by and gaze. If we had universal basic income, they’d say something about it too.

It’s really nice that we only got AIs that draws pictures and bitcoin that helps with stolen money.

Re: Try Stable Diffusion's Img2Img Mode

#62

Earlier quoted context omitted.

I’m not gonna lie, I’m quite disappointed at all the anti-NSFW shenanigans.

I'm not gonna lie, I'm quite disappointed by the wilful obstinance on HN whenever there is a discussion of anti-NSFW or anti-racism filters that are present on today's artificial intelligence research. First, the filters are only on the interface. The research is public. That's why there has been an explosion of new implementations. You are welcome to run the code yourself and make the horniest model you like. But mo…

[deleted]

Re: Try Stable Diffusion's Img2Img Mode

#63
post #12

Earlier quoted context omitted.

"john wayne as captain america" is spot on. It has so many images of both to work with.

i tried "john wayne was a nazi" but was left disappointed

Why were you disappointed? I haven't tried that with img2img, but as just a regular text prompt without any fancy prompt engineering I get results in line with what I'd expect (kind of a Hugo Boss/Cowboy crossover, e.g. this: https://imgur.com/a/gcKi3WG)

Re: Try Stable Diffusion's Img2Img Mode

#64
post #52
post #50

Earlier quoted context omitted.

I'm impressed that running it on CPU only made it ~20x slower. How did you do it?

Nah that's normal. It's why GPUs are the usual thing for AI. Any crap, old, weak gpu with 4gb memory would run circles around a cpu It's often easier to actually get models to run on CPU, due to simpler install configs and more available memory. Just painful to get a result out of it. Which might help keep the install simple, because it's not even worth optimizing

> Any crap, old, weak gpu with 4gb memory would run circles around a cpu

Not really.

https://news.ycombinator.com/item?id=32635086

I’m not even sure it works well - if at all - with 4Gb.

In any case, it’s impressive even if it takes minutes. And it’s not like you need to be there to make it work. You can create a list of prompts, let it do its thing and check the results later.

Re: Try Stable Diffusion's Img2Img Mode

#65
post #13

If you have a GPU with >4GB of VRAM and you want to run this locally, here's a fork of the Stable Diffusion repo with a convenient web UI: https://github.com/hlky/stable-diffusion It supports both txt2img and img2img. (Not affiliated.) Edit: Incidentally, I tried running it on a CPU. It is possible, but it took 3 minutes instead of 10 seconds to produce an image. It also required me to hack up the script in a really…

Just tried it on Ubuntu 22.04. And it's working! Had to install python-is-python3 and conda, but it's up now. Fun, thanks.

How do you set up the model? Instructions only say "Download the model checkpoint. (e.g. from huggingface)", but I can't find instructions there on how to find a ckpt file, nor exactly what file should I look for.

Re: Try Stable Diffusion's Img2Img Mode

#66

Earlier quoted context omitted.

Just tried it on Ubuntu 22.04. And it's working! Had to install python-is-python3 and conda, but it's up now. Fun, thanks.

How do you set up the model? Instructions only say "Download the model checkpoint. (e.g. from huggingface)", but I can't find instructions there on how to find a ckpt file, nor exactly what file should I look for.

You'll need this file: https://huggingface.co/CompVis/stable-diffusion-v-1-4-origin... - Before you can download it you have to accept the T&C at https://huggingface.co/CompVis/stable-diffusion-v-1-4-origin...

Re: Try Stable Diffusion's Img2Img Mode

#67
I've been trying to get some sensible images out of my descriptions, but I fail miserably.

In this case I had the prompt "cow chewing bone" with 4 squares representing the two pair of feet, the body and the head. None cared about chewing on a bone.

With DALL·E 2 I tried to get an image of a little girl building sandcastles and a monster threatening her:

"little scared girl building a sandcastle and a big angry monster is looking at her."

"little scared girl building a sandcastle six damaged sandcastles are to her side. a big angry monster is threatening her. it is dark." https://imgur.com/a/f5FFKOi

"little scared girl building a sandcastle with six damaged sandcastles to her side and a big angry monster threatening her"

Is there some kind of structure the sentences should follow?

Re: Try Stable Diffusion's Img2Img Mode

#68
post #67

I've been trying to get some sensible images out of my descriptions, but I fail miserably. In this case I had the prompt "cow chewing bone" with 4 squares representing the two pair of feet, the body and the head. None cared about chewing on a bone. With DALL·E 2 I tried to get an image of a little girl building sandcastles and a monster threatening her: "little scared girl building a sandcastle and a big angry monste…

Dalle is bad at being instructed to have an exact count of items in the picture. Ask for 6 kittens and you get 7 & each kitten will be much more "wrong" than a piture of a single kitten.

Dalle is is bad at positional prompts. Ask for somethi g to be in the top rightbhand corner and it will appear bottom centre

Re: Try Stable Diffusion's Img2Img Mode

#70
post #48

Earlier quoted context omitted.

There you go. We were promised quite far fetched things like 'flying cars', 'time travel', 'life extension' and 'universal basic income'. Instead we get Deep Learning AI's being used for generating and faking sentences, images, videos, voices, code and digital art all being trained on mountains of data in data centers all significantly contributing to the already burning up of the planet to no benefit and no efficien…

You must be kidding it’s hard labor creating images and you are complaining that it’s going to be easier in the future? You are complaining that the skill gap is eradicated and it now only depends on your ideas to produce compelling work? That’s absolutely insane everything that takes labor intensive tasks and makes them basically free is the future. This is not dystopian it’s our future, a future in which you and yo…

The skill gap was what made it valuable. There is not that much value in flooding everyone with generated images that have no story, no person behind it, no effort, no reason to exist, out of place. “Art” is as much about the final image as it is about the person who made it, why she made it and how she got there.

Also there is pretty dystopian angle because these tools balantly stole all the work of artists by “learning” on their work (without pemission). Calling it learning is too nice we all know its just huge visual pattern copy mashup paste machine. People are not gonna invent with this - they just add name of artist they like to the prompt and get to call that their work.

Its not liberation - its exploitation. Its gonna destroy peoples lives by making their hard earned skills obsolete.

Then again it probably wont be so bad. Artists wont disappear and people wont suddenly become artists just because they can write into prompt.

Post reply on HN