Live data from Hacker News

Stable Diffusion XL 1.0

techcrunch.com

141–150 of 182 posts

Re: Stable Diffusion XL 1.0

#141
post #83

Earlier quoted context omitted.

Any particularly useful resources for looking into prompts?

This AI Horde UI has, IMO, some really good templates and suggestions: https://tinybots.net/artbot

Hey! Creator of ArtBot here. Thanks for plugging the site!

For those not aware, here's an interesting fact about ArtBot (and the AI Horde in general) -- we've been running an A/B test with Stability.ai for the last 3 weeks or so related to SDXL [1].

Any time a user generates an image using SDXL_beta on the AI Horde, they get two images back. They pick which image they think is best for the given prompt. This data is sent back to Stability.ai in order to help improve their image models.

In a similar vein, LAION partnered with the AI Horde earlier this year in order to gather aesthetics ratings for improving various image datasets. [2]

It's a cool little open source community and there's just a ton of stuff going on.

[1] https://dbzer0.com/blog/stable-diffusion-xl-beta-on-the-ai-h...

[2] https://laion.ai/blog/laion-stable-horde/

Re: Stable Diffusion XL 1.0

#142
post #139

Earlier quoted context omitted.

So what if art is devalued? We are hardwired to appreciate beauty so art of some form will always be sought. Obviously there is the matter of artists losing their livelihoods but that is also an inevitable outcome of progress and always has been.

Lol. I hate this industry sometimes. It's not obvious, to me at least, that our society should accept the automation of the production of culture. Maybe I'm a luddite or whatever lazy quip you'd like to use, but I prefer the story of a human mastering a skill and producing something beyond contextless aesthetic sludge.

Yeah, that's why I always laugh when I see people use cars instead of walking 100 miles. It's an inspiring story for people to do marathons, and cars devalue that.

Re: Stable Diffusion XL 1.0

#143

I tried it in dreamstudio. Like all the other image generators I've tried, it's rubbish at drawing a piano keyboard or an accordion. (Those are my tests to see if it understands the geometry of machines.) A couple of accordion pictures do look passable at a distance. Another test: how well does it do at drawing a woman waving a flag? One thing that strikes me is that it generates four images at a time, but there is l…

My go-to test is "elephant riding unicycle". Neither Midjourney nor Stable Diffusion XL is capable of doing this.

Re: Stable Diffusion XL 1.0

#144
post #61

I always wondered why the vision models don't seem to be following the whole "scale up as much as possible" mantra that has defined the language models of the past few years (to the same extent). Even 3.5 billion parameters is absolutely nothing compared to the likes of GPT-3, 3.5, 4, or even the larger open-source language models (e.g. LLaMA-65B). Is it just an engineering challenge that no one has stepped up for ye…

Diffusion is relatively compute intensive compared to transformers llms, and (in current implementation) doesn't quantize as well. A 70B parameter model would be very slow and vram hungry, hence very expensive to run. Also, image generation is more reliant on tooling surrounding the models than pure text prompting. I dont think even a 300B model would get things quite right through text prompting alone.

Hmm this is a good point, diffusion requires several (many?) inference passes as you refine the noise into an image, right? Makes sense that this is more expensive to scale up. Thanks for the explanation!

Re: Stable Diffusion XL 1.0

#145
post #141

Earlier quoted context omitted.

This AI Horde UI has, IMO, some really good templates and suggestions: https://tinybots.net/artbot

Hey! Creator of ArtBot here. Thanks for plugging the site! For those not aware, here's an interesting fact about ArtBot (and the AI Horde in general) -- we've been running an A/B test with Stability.ai for the last 3 weeks or so related to SDXL [1]. Any time a user generates an image using SDXL_beta on the AI Horde, they get two images back. They pick which image they think is best for the given prompt. This data is…

Artbot is an amazing and criminally underappreciated project, I try to find an excuse to plug it wherever I can.

Re: Stable Diffusion XL 1.0

#146
post #73

Is this pre-censored like their other later models?

No. From what I’ve gathered was trained on human anatomy, but not straight up porn. What they tried for 2.0/2.1 was way too overdone, to the point where if I prompted “princess Zelda,” the generation would only look mildly like her. Presumably they just didn’t have many images of people in the training. 1.5 and SDXL both work fine of that front. Fine tuners will quickly take it further, if that’s what you’re after.

I don't think 1.x was trained on porn either?

I seem to remember the issue with 2.x is that they removed all the commercial art from top-notch illustrators from the training data due to the backlash, so it was just way worse at generating great-looking things, which is all the user cares about. So the community stayed on their custom-trained models derived from SD 1.5 (which, yes, often included porn).

Re: Stable Diffusion XL 1.0

#147
post #134

It's often said porn drives technology. I clicked through the links in the article, since they sounded technically interesting. They led to AI-generated porn. Those, in turn, led to pages about training SD to generate porn. Now, two disclaimers: 1) I am not interested in AI-generating porn 2) I haven't followed SD in maybe 6-9 months With those out-of-the-way, the out-of-the-box tools for fine-tuning SD are impressiv…

> the progress seems to be entirely driven by the anime porn community:

Its not entirely driven by porn communities, and the porn communities driving it aren’t entirely anime porn communities (and the anime communities driving it aren’t entirely porn communities.)

But, yeah, the anime + porn/fetish art + furry + rpg art + scifi/fantasy art communities, and particularly the niches in the overlap of two or more of those are, pretty significant.

> If this works for images other than naked and cartoon women

It does, and while it may not be large proportionally compared to the anime-porn stuff, there’s a lot of publicly distributed fine tuned checkpoints, LoRas, etc., demonstrating that it does.

Re: Stable Diffusion XL 1.0

#148
post #134

It's often said porn drives technology. I clicked through the links in the article, since they sounded technically interesting. They led to AI-generated porn. Those, in turn, led to pages about training SD to generate porn. Now, two disclaimers: 1) I am not interested in AI-generating porn 2) I haven't followed SD in maybe 6-9 months With those out-of-the-way, the out-of-the-box tools for fine-tuning SD are impressiv…

It absolutely works for things other than naked and cartoon women. Here are some generations of my daughter and dog (together!). I believe most of these are from a fine tuned model of them and not an extracted LoRA, though I use that sometimes too: https://imgur.com/a/naHgnel

Re: Stable Diffusion XL 1.0

#149

Earlier quoted context omitted.

https://dreamstudio.ai/

What models does dreamstudio use? I couldn't see how to view them without logging in.

Dreamstudio (and ClipDrop, also) uses Stable Diffusion, gettig new SD models generally before public release (both are owned by StabilityAI.)

Re: Stable Diffusion XL 1.0

#150

I am completely uninformed in this space. Would someone be kind to explain what the current state of the art in image generation is (how does this compare to Midjourney and others)? How do open source models stack up? Also what are the most common use cases for image generation?

I don't know what the use case is for other people is, but I've been playing around with book covers. This one took about two weeks, but it was my first real try and I was still learning how. Composition is a little off. The one I'm working on now is going faster (and better).

https://imgur.com/a/CxX5eYj

I've found that I rarely get a usable image completely as-is. It might take 5 or 10 generations to find something sort of ok, and even then I end up erasing the bad parts and letting it in-paint (which again takes multiple attempts). The T-rex had like 7 legs and two jaws, but was otherwise close to what I wanted... just keep erasing extra body parts until the in-painter finally takes a hint.

I was also going to do a few book covers for some Babylon 5 books, but it does so bad on celebrity faces. Looked like Koenig's mutant love child with Ernest Borgnine. Dunno what to do about that. I keep wondering if I shouldn't spend the next 10 years putting together my own training set of fantasy and science fiction art.

Post reply on HN