Live data from Hacker News

SDXL Turbo: A Real-Time Text-to-Image Generation Model

stability.ai

151–157 of 157 posts

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#151

Earlier quoted context omitted.

Is that the correct link? I've never heard of A40s, the link is to release notes from a year and two months ago, and SD XL just came out a month or two ago. Hard for me to get to "SD XL cannot [be finetuned effectively]" from there.

A40: https://www.techpowerup.com/gpu-specs/a40-pcie.c3700 I have not heard about the team upgrading or downgrading from the hardware mentioned there, so I assumed it's still the same hardware they use. >SD XL just came out a month or two ago About 4.5 months actually. For the SDXL cannot be finetuned efficiently claim, an attempt at a finetune was released here: https://huggingface.co/hakurei/waifu-diffusion-xl The t…

> The team was given early access by StabilityAI to SDXL0.9 for this.

SDXL0.9 may have been a worse starting point than SDXL1.0, and certainly there's been more time put in and experience developed finetuning SDXL in the time since WD trained against SDXL0.9 by the people who have released the huge pile of finetunes that have been released since.

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#152

Earlier quoted context omitted.

Are there any good resources for learning a lot of "syntactic sugar" terms? This is new to me, but I'd love to know more.

It is completely dependent on the model. Civit dot ai has model showcases as well as fine-tune showcases, and you can click any image or press the (i) to see the generation info. Some models like natural language prompts - "draw me a pterodactyl tanning at a beach", some prefer shorthand (danbooru style clip) - "1man, professor, classroom, chalkboard, white_hair, suit", and some work with a mixture of the above as we…

> Civit dot ai

The site you are thinking of is https://civitai.com/ not "civit dot ai".

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#153

the clipdrop demo doesn't inspire much confidence with how bad the generations are, we're talking 2021 levels, not to mention everything is NSFW somehow, should probably work on those filters.

I tested a bit and the quality for photorealistic images is surprisingly bad, and definitely worse than LCM and of course normal SDXL. For more artistic images, SDXL Turbo fares better. Unlike normal SDXL, you're required here to use the old-fashioned syntatic sugar like "8k hd" and "hyperrealistic" to align things.

In my own testing (using ComfyUI), the best of the "fast gen" techniques for sdxl is using the Turbo model [0], but using the LCM sampler with the sgm_uniform scheduler (which is normal for LCM) with it, and running it up to 4-10 iterations instead of just one. I think StabilityAI demos are using Euler A with the normal base scheduler, and running a single iteration (which is cool for a max-speed demo, and its awesome for that speed, but its leaving a lot of quality on the table that you can get with a few more iterations especially with the LCM/sgm_uniform sampler/scheduler combo.) Bumping CFG up slightly helps, too (but I think adds another performance hit, because I think the demos are running at CFG 1, which AIUI disables CFG and reduces computations per iteration.)

> Unlike normal SDXL, you're required here to use the old-fashioned syntatic sugar like "8k hd" and "hyperrealistic" to align things.

That's not "syntactic sugar", and its not particular my experience that it is needed with sdxl turbo.

[0] actually, differencing the base sdxl model from the turbo model to get a "turbo modifier", and then combining that with a good SDXL-based checkpoint, because StabilityAI's base models are pretty ho-hum compared to decent community checkpoints derived from them, but that is kind of a peripheral issue.

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#154

SDXL is already very very slow when compared to SD 1.5. They are claiming 200ms for 512x512 image in SDXL on A100. We need SD 1.5 turbo for even faster generation.

Well, its 2.1 instead of 1.5, but:

https://huggingface.co/stabilityai/sd-turbo

I assume that this was ready for release because Stable Video Diffusion (which is also an SD2.1-based model) is essentially this plus a motion model.)

I wouldn't be surprised if their hosted-only SD1.6 beta is, or has, a turbo version, and if that gets released publicly, that's where we'll see an SD1.x turbo.

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#155
post #38

SDXL is already very very slow when compared to SD 1.5. They are claiming 200ms for 512x512 image in SDXL on A100. We need SD 1.5 turbo for even faster generation.

is it known how much larger SDXL is compared to SD1 or SD2?

SD1.x is around 1.1B parameters including VAE, SD2.x is slightly more (uses the same UNet and VAE, but a bigger text encoder; not finding stats as quickly as I'd like), and SDXL is 3.5B parameters (single model only, but StabilityAI's preferred base + refiner model setup is effectively 6.6B parameters -- some bits are shared.)

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#156

Earlier quoted context omitted.

What do you mean by extreme requirements? There's lots of SDXL fine tunings available at civit, like https://civitai.com/models/119012/bluepencil-xl for anime. The relevant discords for models/apps are full of people doing this at home. Or are you looking at some very specific definition / threshold for fine tuning here?

There is something I find rather hard to communicate about the difference between these models on civitai and what I think a competent model should be able to do. I'd describe them like "a bike with no handlebars" because they are incredibly difficult to steer to where you want. For example if you look at the preview images like this one: https://civitai.com/images/3615715 The model seems to have completely ignored a…

> the model seems to have completely ignored a good 35% of the text input, Well, when you blindly cargo cult a prompting style designed to work around issues with SD1.x in SDXL and in the process spam several hundred tokens, mostly slight variants, into the 75-ish-token window (which, yes, the UIs use a merging strategy to try to accommodate), you have that problem.

> most egregiously I find the (flat chest:2.0)

The flat chest is fighting to compensate the also heavily weighted (hands on breasts:1.5) which not only affects hand placement but also the concept of "breasts", and the biases trained into many of the community models with that term mean that having that concept in the prompt and heavily weighted takes a lot to counteract. So, no, I don't think its ignoring that.

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#157

Earlier quoted context omitted.

>do laws adequately protect people (kids)? Can they? Will this force a shift towards actually chasing producers, distributors, and diddlers? It's extremely complicated. Actual CSAM is very illegal, and for good reason. However, artistic depictions of such are... protected 1st Amendment expression[0]. So there's an argument - and I really hate that I'm even saying this - that AI generated CSAM is not prosecutable, as…

Could this perhaps fall under something like trademark, like an unauthorized use of self, I'm sure I've heard of some celebrity cases that were for similar.

You're probably thinking of the Right of Publicity laws some US states have.
Post reply on HN