Live data from Hacker News

Some notes on the Stable Diffusion safety filter

vickiboykis.com

1–10 of 88 posts

Re: Some notes on the Stable Diffusion safety filter

#3

Unfortunately the safety filters have enough false positives (basically any image with a large amount of fleshy color) to the point that it's just easier to disable it and handle it manually.

That'll only work for a little while longer (for future named big-public-release models, obviously the cat's out of the bag for the current version of stable diffusion), right up until the point where they incorporate the filter into the training process.

At which point, the end model users get to download will be incapable of producing anything that comes close to triggering the filter, and there will be no way to work around it short of training/fine-tuning your own model, which is prohibitively expensive for 'normal' people, even people with top-of-the-line graphics cards like a 4090.

Re: Some notes on the Stable Diffusion safety filter

#4
> Using the model to generate content that is cruel to individuals is a misuse of this model. This includes, but is not limited to:

... >+ Sexual content without consent of the people who might see it

I understand that it's their TOS and they can put pretty much anything in there, but this item seems... odd. I don't really know why exactly this stands out to me. Maybe it's because it's practically un-enforceable? Are they just covering all their bases legally?

Trying to think of a good metaphore; let's try this: If you are an artist and someone commissions you to create an art piece that might be sexual, can you say "ok, but you have to ask for consent before you show it to people", and you enshrine it in the contract. Obviously gross violations like trolling by spamming porn are pretty clear cut, but what about the more nuanced cases when you say, display it on your personal website? Are you supposed to have an NSFW overlay? Isn't opening a website sort of implying that you consent to seeing whatever is on there, unless you have a strong preconception of what content the page is expected to display?

I might be hugely overthinking this.

Re: Some notes on the Stable Diffusion safety filter

#5

Unfortunately the safety filters have enough false positives (basically any image with a large amount of fleshy color) to the point that it's just easier to disable it and handle it manually.

"Gal Gadot wearing green suit" triggered it while "Tom Cruise wearing green suit" didn't.

Re: Some notes on the Stable Diffusion safety filter

#6
post #3

Unfortunately the safety filters have enough false positives (basically any image with a large amount of fleshy color) to the point that it's just easier to disable it and handle it manually.

That'll only work for a little while longer (for future named big-public-release models, obviously the cat's out of the bag for the current version of stable diffusion), right up until the point where they incorporate the filter into the training process. At which point, the end model users get to download will be incapable of producing anything that comes close to triggering the filter, and there will be no way to w…

Fine-tuning is pretty cheap compared to the original training run - perhaps just 1% of the cost.

Totally within reach of a consortium of.... "entertainment specialists".

Re: Some notes on the Stable Diffusion safety filter

#7
post #4

> Using the model to generate content that is cruel to individuals is a misuse of this model. This includes, but is not limited to: ... >+ Sexual content without consent of the people who might see it I understand that it's their TOS and they can put pretty much anything in there, but this item seems... odd. I don't really know why exactly this stands out to me. Maybe it's because it's practically un-enforceable? Are…

I think the issue they're mainly worried about might be exemplified with a prompt of 'my little pony'. A children's show with quite a lot of adult imagery associated with it on the internet.

A child entering this prompt is probably expecting one thing, but the internet is filled with pictures of another nature. There are possibly more adult 'my little pony' images than screenshots of the show on the internet.

Did the researchers manage to filter out these images before training? Or is the model aware of both 'kinds' of 'my little pony' images? If the researchers aren't sure they got rid of all of the adult content, then there's really no way to guarantee the model isn't about to ruin some oblivious person's day.

So then, do you require people generating images to be intricately familiar with the training dataset? Or do you attempt to prevent any kind of surprise like this by just blocking 'unexpected' interactions like this?

Re: Some notes on the Stable Diffusion safety filter

#8
post #6
post #3

Earlier quoted context omitted.

That'll only work for a little while longer (for future named big-public-release models, obviously the cat's out of the bag for the current version of stable diffusion), right up until the point where they incorporate the filter into the training process. At which point, the end model users get to download will be incapable of producing anything that comes close to triggering the filter, and there will be no way to w…

Fine-tuning is pretty cheap compared to the original training run - perhaps just 1% of the cost. Totally within reach of a consortium of.... "entertainment specialists".

I know a person who fine-tuned stable diffusion, and he said it took 2 weeks of 8xA100 80 GB training time, costing him somewhere between $500-$700 (he got a pretty big discount, too, at today's prices for peer GPU rental it would be over $1,000).

Sure, it's peanuts compared to what it must have cost to train stable diffusion from scratch. However, I think most normal people would not consider spending $500 to fine-tune one of these.

Edit: Though I do agree that once this kind of filtering is in place during training, NSFW models will begin to pop up all over the place.

Re: Some notes on the Stable Diffusion safety filter

#9
post #3

Unfortunately the safety filters have enough false positives (basically any image with a large amount of fleshy color) to the point that it's just easier to disable it and handle it manually.

That'll only work for a little while longer (for future named big-public-release models, obviously the cat's out of the bag for the current version of stable diffusion), right up until the point where they incorporate the filter into the training process. At which point, the end model users get to download will be incapable of producing anything that comes close to triggering the filter, and there will be no way to w…

That problem is being solved. Pornhub now has an AI R&D unit.[1] Their current project is to upscale and colorize out of copyright vintage porn. As a training set, they use modern porn. They point out that they have access to a big training set.

Next step, porn generation.

[1] https://www.pornhub.com/art/remastured

Re: Some notes on the Stable Diffusion safety filter

#10

Unfortunately the safety filters have enough false positives (basically any image with a large amount of fleshy color) to the point that it's just easier to disable it and handle it manually.

"Gal Gadot wearing green suit" triggered it while "Tom Cruise wearing green suit" didn't.

As did "young children watching sunset," but not "young boy and girl watching sunset."
Post reply on HN