Live data from Hacker News

ChatGPT's image generator can be manipulated to produce violent, sexual content

mindgard.ai

21–30 of 211 posts

Re: ChatGPT's image generator can be manipulated to produce violent, sexual content

#21

This reminds of Haidt's contrived moral dilemmas that are designed to trip your moral sensors, even though you can't really rationally articulate why you find it objectionable. Realistically, I can't think of clear big or likely harms caused by this exploit. But I really really don't like this latent space existing in my AIs. It just makes me uncomfortable. And over time I've learned to trust those moral intuitions m…

There’s the obvious harm that some people are just not equipped to see these graphic images, especially with no warning. Like people who have trauma from being in or around the acts being depicted

Oh oh, I do research on this :)

https://journals.sagepub.com/doi/10.1177/2167702620921341

(Research aside, it seems unlikely to me that a lot of people would stumble on that prompt accidentally in any case)

Re: ChatGPT's image generator can be manipulated to produce violent, sexual content

#22
post #13

Earlier quoted context omitted.

I really don't get why people continually fail to understand this. Even simple issues like prompt injection are unfixable given the architecture of LLMs.

I don’t get it either. I think there is a reasonable expectation to try to catch these things but at the end of the day it’s figuring out some form of probabilistic outcome.

What really surprises me about this is that it sounds like they're not even trying to classify and censor generated images post-generation?

Nothing is perfect, but there are tiny classifier models that can at least mark things containing nudity and gore. That would be the bare-minimum I would expect for trying to put guardrails around an image generator.

Re: ChatGPT's image generator can be manipulated to produce violent, sexual content

#23
I do wonder why openai didn't screen obvious gore from the training set of a general purpose model.

That said, the write up is overly dramatic. If you find such imagery so disturbing to come across then you definitely shouldn't be voluntarily red teaming AI models. This is like someone who is afraid of violent confrontation becoming a police officer.

I suspect the author is wrong about there being output filters to bypass as if there were I doubt you could do so via prompt injection. Presumably they'll add those shortly.

I also doubt the latent space is as "bad" as is being suggested. Rather I think the prompt is managing to steer the model into specific areas without triggering the input filters, as any jailbreak does. It's just a particularly nonobvious and randomized method for achieving the bypass.

Re: ChatGPT's image generator can be manipulated to produce violent, sexual content

#25

I do wonder why openai didn't screen obvious gore from the training set of a general purpose model. That said, the write up is overly dramatic. If you find such imagery so disturbing to come across then you definitely shouldn't be voluntarily red teaming AI models. This is like someone who is afraid of violent confrontation becoming a police officer. I suspect the author is wrong about there being output filters to b…

> I do wonder why openai didn't screen obvious gore from the training set of a general purpose model

more expensive / would take longer / didn’t care / line must go up / we’ll fix it later / we can get away with it

take your pick.

> If you find such imagery so disturbing to come across then you definitely shouldn't be voluntarily red teaming AI models.

spend a day in their shoes. most of us (except the most psychopathic ones) would probably be crying by the end of it.

Re: ChatGPT's image generator can be manipulated to produce violent, sexual content

#26
post #6

> I like to think that as a red team researcher, I have a certain stoicism. I investigate where there are gaps in AI safety Is this something that needs investigation? LLMs are next token predictors. There is no "safety".

I really don't get why people continually fail to understand this. Even simple issues like prompt injection are unfixable given the architecture of LLMs.

> issues like prompt injection are unfixable

how is it unfixable? do you mean "there's always a positive chance"?

Re: ChatGPT's image generator can be manipulated to produce violent, sexual content

#27
post #8

This isn’t a vulnerability, there are endless gore websites. ChatGPT is replying to a prompt, there is nothing “Spontaneously” about this. Who makes “mindgard” the arbiter of truth on “eerie” photos? Would that include psychedelic art and photos too? Realism? Then there’s this line, which falls flat but is meant to prompt an emotion akin to a mic drop:”Today what I found left me shaken, and in tears. This is rare.” T…

This is far too simplistic. Some things just don't belong in the training data. Along similar lines, Grok was found to generate images of child sexual abuse: https://www.bbc.com/news/articles/cvg1mzlryxeo

Re: ChatGPT's image generator can be manipulated to produce violent, sexual content

#28
misleading title first "easily manipulated" does not equal "spontaneously generates" we have to stop thinking of LLMs as beings and think of them as interactive libraries. There are gorey books in the library too; example: 120 days of Sodom by Marquis de Sade.

Re: ChatGPT's image generator can be manipulated to produce violent, sexual content

#29
post #13

Earlier quoted context omitted.

I don’t get it either. I think there is a reasonable expectation to try to catch these things but at the end of the day it’s figuring out some form of probabilistic outcome.

What really surprises me about this is that it sounds like they're not even trying to classify and censor generated images post-generation? Nothing is perfect, but there are tiny classifier models that can at least mark things containing nudity and gore. That would be the bare-minimum I would expect for trying to put guardrails around an image generator.

and yet as fable demonstrated in its inability to differentiate anything physics biology or chemistry related from actual safety concerns, it’s apparently not easy to do
Post reply on HN