Live data from Hacker News

Some notes on the Stable Diffusion safety filter

vickiboykis.com

61–70 of 88 posts

Re: Some notes on the Stable Diffusion safety filter

#61
post #45

Earlier quoted context omitted.

>not realistic The odds are zero. 1/2^256 = 0. In cryptography these odds are treated as zero until you generate close to 2^128 images. Unfortunately there's no word in natural English to describe how unlikely. The most precise is "zero".

You should spend some time with an internet search engine and the term "perceptual hashing". What you're talking about is another type of hashing, which can be useful for classifying image files , but not images . The former has a very concrete definition that is specified down to the bit; the latter is a fuzzy space because it's trying to yield similar (not necessarily identical) hashes for images that humans consid…

Oh wow https://www.apple.com/child-safety/pdf/CSAM_Detection_Techni... so they essentially just use CNN output to automatically determine whether to report people to the authorities? For some reason I assumed they were just comparing the files they knew to be CSAM.

Yeah that's bad. What about deepdream/CNN reversing? Couldn't a rogue apple engineer just create a innocuous looking false positive, say a cat picture, share it on Reddit, and everybody who downloads it is flagged to police for CSAM?

Re: Some notes on the Stable Diffusion safety filter

#62
post #45

Earlier quoted context omitted.

You should spend some time with an internet search engine and the term "perceptual hashing". What you're talking about is another type of hashing, which can be useful for classifying image files , but not images . The former has a very concrete definition that is specified down to the bit; the latter is a fuzzy space because it's trying to yield similar (not necessarily identical) hashes for images that humans consid…

Oh wow https://www.apple.com/child-safety/pdf/CSAM_Detection_Techni... so they essentially just use CNN output to automatically determine whether to report people to the authorities? For some reason I assumed they were just comparing the files they knew to be CSAM. Yeah that's bad. What about deepdream/CNN reversing? Couldn't a rogue apple engineer just create a innocuous looking false positive, say a cat picture, sh…

The CSAM flagging generally isn’t reported to police to prevent the situation you describe. Google would get the report and once some threshold is reached, a person reviews the report(s) and decides if the police are notified.

Re: Some notes on the Stable Diffusion safety filter

#63

Earlier quoted context omitted.

A child entering this prompt is probably expecting one thing, but the internet is filled with pictures of another nature. There are possibly more adult 'my little pony' images than screenshots of the show on the internet. So everyone has to have gimpy AI just because parents can't be expected to take responsibility for what their child does and does not see? Why the fuck is a child being allowed to play with somethin…

Just to be clear, the child was just an example of someone who could theoretically experience 'cruel' treatment from the current version of stable diffusion. I'm absolutely not recommending people let their children use the model unsupervised. It doesn't have to be a parenting problem, though. The same could be said (for example) of a random mother trying to get inspiration for a 'my little pony' birthday cake for th…

Interestingly, better results might be achieved by exposing the model to a large corpus of apropriately tagged NSFW data, so that the prompt may exclude it. I imagine img2img could also make an image SFW, or vice-versa. I'd be curious to know what kind of alterations it would make.

Re: Some notes on the Stable Diffusion safety filter

#64
post #45

Earlier quoted context omitted.

You should spend some time with an internet search engine and the term "perceptual hashing". What you're talking about is another type of hashing, which can be useful for classifying image files , but not images . The former has a very concrete definition that is specified down to the bit; the latter is a fuzzy space because it's trying to yield similar (not necessarily identical) hashes for images that humans consid…

Oh wow https://www.apple.com/child-safety/pdf/CSAM_Detection_Techni... so they essentially just use CNN output to automatically determine whether to report people to the authorities? For some reason I assumed they were just comparing the files they knew to be CSAM. Yeah that's bad. What about deepdream/CNN reversing? Couldn't a rogue apple engineer just create a innocuous looking false positive, say a cat picture, sh…

No, there are two hashes used in the Apple system, one public and neural and one hidden, the intent of both is to match specific known images and not unknown new ones, and the result of passing both hashes is a manual review and not automatic reporting. I've never seen a published attack that would actually be a problem; they all misread how the system worked.

(Also, it's not reported to the police but to NCMEC, which is not a government agency. This is for 4th amendment privacy reasons.)

Re: Some notes on the Stable Diffusion safety filter

#65

I can understand giving a user the option to filter out something they might not want to see. But the idea that the technology itself should be limited based on the subjective tastes and whims of the day makes my stomach churn. It's not too disconnected from altering a child's brain so that he is incapable of understanding concepts his parents don't like.

Nope, it's not like that in any way at all. AI aren't children, they're a technological artefact a group of people assembled, with no more moral dimension than a toaster. The people who would be the targets of Stable Diffusion-generated porn, depictions they did not consent to, they're actual people who's privacy would be harmed. Technology is an artefact of the society that produced it, and is shaped by it's values. This is not the symptom of any sort of problem.

This type of technological fetishism that holds that technology should be developed for it's own sake and that the well being of society is secondary should be discarded. Technology should be developed to make people's lives better or to expand our understanding, not just because it can be. That's how we end up with the proliferation of harmful technologies of no benefit.

Re: Some notes on the Stable Diffusion safety filter

#66
post #4

> Using the model to generate content that is cruel to individuals is a misuse of this model. This includes, but is not limited to: ... >+ Sexual content without consent of the people who might see it I understand that it's their TOS and they can put pretty much anything in there, but this item seems... odd. I don't really know why exactly this stands out to me. Maybe it's because it's practically un-enforceable? Are…

They dont want it being used for porn. I wouldnt either. But they should just say that.

Re: Some notes on the Stable Diffusion safety filter

#67
post #55

Earlier quoted context omitted.

"Gal Gadot wearing green suit" triggered it while "Tom Cruise wearing green suit" didn't.

Might be the word "gal" (which can mean girl or young woman).

Adult porn is filtered too.

Re: Some notes on the Stable Diffusion safety filter

#68

I can understand giving a user the option to filter out something they might not want to see. But the idea that the technology itself should be limited based on the subjective tastes and whims of the day makes my stomach churn. It's not too disconnected from altering a child's brain so that he is incapable of understanding concepts his parents don't like.

> But the idea that the technology itself should be limited based on the subjective tastes and whims of the day makes my stomach churn.

That's your prerogative when you create something. And while I agree with filtering out porn, you can take solice that people will bypass it.

Re: Some notes on the Stable Diffusion safety filter

#69
post #3

Unfortunately the safety filters have enough false positives (basically any image with a large amount of fleshy color) to the point that it's just easier to disable it and handle it manually.

That'll only work for a little while longer (for future named big-public-release models, obviously the cat's out of the bag for the current version of stable diffusion), right up until the point where they incorporate the filter into the training process. At which point, the end model users get to download will be incapable of producing anything that comes close to triggering the filter, and there will be no way to w…

This (training a model with no NSFW content) would be preferable to me. No false positives to worry about. People who do want to generate NSFW stuff can fine-tune or train their own model, nobody owes that functionality to them in freely available ones.
Post reply on HN