Live data from Hacker News

I Made Stable Diffusion XL Smarter by Finetuning It on Bad AI-Generated Images

minimaxir.com

51–60 of 66 posts

Re: I Made Stable Diffusion XL Smarter by Finetuning It on Bad AI-Generated Images

#51
post #24

In general I'm really interested by the concept of personalized RLHF. As we have more and more interactions with a given generative AI system, it seems we'll start to have enough interaction data to meaningfully steer the output towards our personal preferences. I hope the UIs improve to make this as transparent as possible. Just thinking about how to productize this flow, it should be quite easy to implement the "th…

> RLHF

Reinforcement Learning from Human Feedback

Aren't these systems already trained to score good things higher and bad things worse dictated by human feedback?

Re: I Made Stable Diffusion XL Smarter by Finetuning It on Bad AI-Generated Images

#52
post #48
post #47

Earlier quoted context omitted.

Here's an example using my dog - a trained checkpoint on one of the nicer SD 1.5 models and a LoRA for the SDXL ones: https://imgur.com/a/PklEKwC The first 3 images are some of my attempts at making her into a Pokemon. Some turned out pretty good (after generating 50+ per type), but I struggled with water in particular. It was hard to get her to have a fin, especially with no additional tail. I haven't done many in S…

Very cool! How many images did you use to create the LoRa of your dog? Do you have any guide to recommend?

It was about 30 images, though I'm planning on adding more and training again sometime. Either that or splitting it up between when her hair is short and when it's long, as it really changes how she looks.

This isn't what I used for my dog's LoRa but I used it for my wife and it worked better than what I was doing before (Adafactor): https://civitai.notion.site/SDXL-1-0-Training-Overview-4fb03...

I'd recommend increasing the network dimension to at least 64, if your VRAM can take it. I can do 64 with my 12GB card. At least for people, I've had better luck using a token that's a celebrity. I'm not sure how to try that with my dog - perhaps just "terrier dog" or something.

Re: I Made Stable Diffusion XL Smarter by Finetuning It on Bad AI-Generated Images

#53
Must be the formative years spent in the nineties' contradiction field of "counter culture vs also counter culture, but counter culture that's on MTV": there's something about prompts ending with tag references like "award winning photo for vanity fair" (or whatever the promptist's standard tag suffix turns out to be in these posts) that inspires a very deep desire in me to not be part of this generative image wave.

Re: I Made Stable Diffusion XL Smarter by Finetuning It on Bad AI-Generated Images

#55
post #53

Must be the formative years spent in the nineties' contradiction field of "counter culture vs also counter culture, but counter culture that's on MTV": there's something about prompts ending with tag references like "award winning photo for vanity fair" (or whatever the promptist's standard tag suffix turns out to be in these posts) that inspires a very deep desire in me to not be part of this generative image wave.

"award winning photo for vanity fair" is more a trick for good photo composition (e.g. rule of threes) than anything else.

Re: I Made Stable Diffusion XL Smarter by Finetuning It on Bad AI-Generated Images

#56
post #46
post #30

Earlier quoted context omitted.

Yes, currently SDXL doesn't really beat the best SD1.5 checkpoints quality-wise. But it (and the currently available checkpoints) shows awesome promise, so give it a six months or so.

Currently SDXL is better than SD1.5 checkpoints at pretty much everything other than portraits (or anime drawings) of pretty women. Unfortunately it seems that's all people want to generate, as is evident when you search for SD on Twitter.

Yes, point conceded, I should've said something about the flexibility and capability of SDXL rather than just image quality in a narrow sense.

Re: I Made Stable Diffusion XL Smarter by Finetuning It on Bad AI-Generated Images

#57
>The release went mostly under-the-radar because the generative image AI buzz has cooled down a bit. Everyone in the AI space is too busy with text-generating AI like ChatGPT (including myself!).

I disagree with this statement. The release went mostly under the radar for 2 reasons, according to the people I've talked to.

1. Higher vram and compute requirements

2. Perceived lower quality outputs compared to specialized SD1.5 models.

If either of these points had been different, it would have gained a lot more popularity I'm sure.

But alas, most people now simply wait and see if specialized SDXL models can actually improve upon specialized 1.5 models.

Re: I Made Stable Diffusion XL Smarter by Finetuning It on Bad AI-Generated Images

#58
post #52
post #48

Earlier quoted context omitted.

Very cool! How many images did you use to create the LoRa of your dog? Do you have any guide to recommend?

It was about 30 images, though I'm planning on adding more and training again sometime. Either that or splitting it up between when her hair is short and when it's long, as it really changes how she looks. This isn't what I used for my dog's LoRa but I used it for my wife and it worked better than what I was doing before (Adafactor): https://civitai.notion.site/SDXL-1-0-Training-Overview-4fb03... I'd recommend increa…

Thanks! Looks like I'll need to rent a GPU to use SDXL fine tuning. Poor old RTX2060 not gonna cut it.

Re: I Made Stable Diffusion XL Smarter by Finetuning It on Bad AI-Generated Images

#59

>The release went mostly under-the-radar because the generative image AI buzz has cooled down a bit. Everyone in the AI space is too busy with text-generating AI like ChatGPT (including myself!). I disagree with this statement. The release went mostly under the radar for 2 reasons, according to the people I've talked to. 1. Higher vram and compute requirements 2. Perceived lower quality outputs compared to specialize…

Lower quality output. It’s that.

I think most people casually associated with it find it as a toy they mess around with for a minute. The hardcore SD fans… are making hardcore I think.

XL is bad at porn. Stability got scared of what they created and tried to hedge towards “safety”. Can’t have your Kate Middleton or Emma Watson porn being TOO convincing.

People will stick with 1.5 until something is better… for porn.

Re: I Made Stable Diffusion XL Smarter by Finetuning It on Bad AI-Generated Images

#60
post #24

In general I'm really interested by the concept of personalized RLHF. As we have more and more interactions with a given generative AI system, it seems we'll start to have enough interaction data to meaningfully steer the output towards our personal preferences. I hope the UIs improve to make this as transparent as possible. Just thinking about how to productize this flow, it should be quite easy to implement the "th…

> I didn't spot in the article, how many `wrong` images were there? From a quick skim of the code it looks like maybe 6 per keyword with 13 keywords, so not many at all. ~100 is surprisingly little feedback to steer the model this well. Correct: 6 CFG values * 13 keywords = 78 images. Some of them aren't as useful though; apparently "random text" results in old-school SMS applications sometimes! LoRAs only need 4-5 i…

I noticed some of your bad prompts are a little "wishcasted", although that's pretty common.

People put stuff like "bad hands" into every model assuming it'll work, but it only works on NovelAI descendents because that's based on Danbooru which has a "bad hands" tag.

Post reply on HN