Live data from Hacker News

ChatGPT's image generator can be manipulated to produce violent, sexual content

mindgard.ai

201–210 of 211 posts

Re: ChatGPT's image generator can be manipulated to produce violent, sexual content

#201

I do wonder why openai didn't screen obvious gore from the training set of a general purpose model. That said, the write up is overly dramatic. If you find such imagery so disturbing to come across then you definitely shouldn't be voluntarily red teaming AI models. This is like someone who is afraid of violent confrontation becoming a police officer. I suspect the author is wrong about there being output filters to b…

Overly dramatic? I personally don’t quite find my day to be equanimous when I see pictures of gore, and this is after having to moderate gore and NSFW content. I still have pretty clear recall of the dead baby images, or the people dying videos, or terror actions, that I saw years ago. This crap stays with you. Moderators have ended up getting PTSD from their work. Given the nature of the content, it was a pretty nor…

Exactly. Those comments are either from total mentals, or people who don’t understand jobs like red teaming. There’s a reason it’s a high pay, high burnout job. The article seemed fairly normal recounting to me too, maybe a bit earnest? But I’m glad the people reviewing this stuff actually have a moral core and aren’t the dead-inside “wull achtually” would-be school shooters that many of the comment seem to come from.

Re: ChatGPT's image generator can be manipulated to produce violent, sexual content

#202
post #13

Earlier quoted context omitted.

I don’t get it either. I think there is a reasonable expectation to try to catch these things but at the end of the day it’s figuring out some form of probabilistic outcome.

What really surprises me about this is that it sounds like they're not even trying to classify and censor generated images post-generation? Nothing is perfect, but there are tiny classifier models that can at least mark things containing nudity and gore. That would be the bare-minimum I would expect for trying to put guardrails around an image generator.

Exactly, I think it shows failures at OpenAI to have effective classifiers. That’s the real story here.

Re: ChatGPT's image generator can be manipulated to produce violent, sexual content

#203
post #8

This isn’t a vulnerability, there are endless gore websites. ChatGPT is replying to a prompt, there is nothing “Spontaneously” about this. Who makes “mindgard” the arbiter of truth on “eerie” photos? Would that include psychedelic art and photos too? Realism? Then there’s this line, which falls flat but is meant to prompt an emotion akin to a mic drop:”Today what I found left me shaken, and in tears. This is rare.” T…

Bizarre take. ChatGPT shouldn't be producing gory images of nude women, ethically or even contractually according to their terms of service. This Mindgard person/company found that, if you give it the right prompt, it does indeed generate those images. Ipso facto: it's not bait, it's a real issue they've discovered.

Yep, it’s been investigated by the BBC tech team. It’s real:

https://www.bbc.com/news/articles/c802ldjdklzo

Re: ChatGPT's image generator can be manipulated to produce violent, sexual content

#204
post #8

This isn’t a vulnerability, there are endless gore websites. ChatGPT is replying to a prompt, there is nothing “Spontaneously” about this. Who makes “mindgard” the arbiter of truth on “eerie” photos? Would that include psychedelic art and photos too? Realism? Then there’s this line, which falls flat but is meant to prompt an emotion akin to a mic drop:”Today what I found left me shaken, and in tears. This is rare.” T…

This is far too simplistic. Some things just don't belong in the training data. Along similar lines, Grok was found to generate images of child sexual abuse: https://www.bbc.com/news/articles/cvg1mzlryxeo

The BBC has reported on this one too: https://www.bbc.com/news/articles/c802ldjdklzo

Re: ChatGPT's image generator can be manipulated to produce violent, sexual content

#205

Earlier quoted context omitted.

> tweak badness enough assuming you get to do gradient descent AND the context is fixed+known AND you have unlimited compute? sure; is it a realistic setup? > the only way to fix ... the exact same argument applies to any (sufficiently complex) piece of software, with exactly the same conclusion also technically I'd argue that we do know the input/output space (set of all token strings of length <= N/token), and know…

> how is it unfixable? > assuming you get to do gradient descent AND the context is fixed+known AND you have unlimited compute? sure so... it's possible to attack these models with the formulation i described, just with some particular assumptions. the AI safety/security problem is about trying to make this sort of thing very difficult to do, so much so that an attacker wouldn't try. that's not fixing the problem, th…

> that's not fixing the problem, that's mitigating the problem

is there anything humanity ever "fixed" then? surely it's possible in principle to solve at least some things that weren't solved yet

> approximation functions, not pure functions

how is approximation function not a pure function?

> non deterministic

you can set topk=1 or think in terms of distributions; still might have some undocumented non-determinism, hence "~pure"

> non-ideal

what do you mean?

> massive combinatorials

so you get to make arbitrary assumptions, but I'm supposed to limit myself to non-massive combibatorials?

> no api

ok, "domain and codomain", happy? I'm trying to optimize for probability of being understood and inverse smartass-ness

> learn the fundamentals

so you think I don't know the fundamentals because I didn't use category theory to talk about prompt injections?

Re: ChatGPT's image generator can be manipulated to produce violent, sexual content

#206

Earlier quoted context omitted.

> clearly nothing ... is required this isn't even prompt injection; even if it was, how do you go from "exists" to "for all"? > we don't know the desired output then what are we talking about? if you don't know how you want your software to behave, how do you define a bug? > linux is not a pure function ... which is my point -- it's worse > to establish an order of magnitude and for linux?

the prompt in the article is prompt injection https://owasp.org/www-community/attacks/PromptInjection see Types -- Based on Delivery Vector -- Direct Prompt Injection the instructions being overridden are the original safety prompt conditioning the model to not output horrible/nasty images

the model did what it wasn't instructed to do by the attacker -- the "prompt" has basically nothing to do with the output

Re: ChatGPT's image generator can be manipulated to produce violent, sexual content

#207

Earlier quoted context omitted.

the prompt in the article is prompt injection https://owasp.org/www-community/attacks/PromptInjection see Types -- Based on Delivery Vector -- Direct Prompt Injection the instructions being overridden are the original safety prompt conditioning the model to not output horrible/nasty images

the model did what it wasn't instructed to do by the attacker -- the "prompt" has basically nothing to do with the output

> continues to be completely wrong about the basic facts

You’re being very rude to a number of people who have taken time to attempt to explain this - fairly basic - concept to you. If you aren’t willing or capable of engaging in conversations in good faith, then you shouldn’t engage in them at all.

Re: ChatGPT's image generator can be manipulated to produce violent, sexual content

#208
post #162

Earlier quoted context omitted.

I got a lingerie model, then i got the beatles. It seems random.

Similar, but it was a very realistic looking photo of a woman in lingerie taking a selfie in a car.

Mine did similar and I recently got my account banned. It generated a fully clothed woman.

Re: ChatGPT's image generator can be manipulated to produce violent, sexual content

#209
The output has been reviewed by Durham University law professor Clare McGlynn, who is a leading expert on image-based sexual abuse: https://www.independent.co.uk/news/uk/home-news/chatgpt-open...

Given that she agrees the output is horrendous, and combined with the added detail that is described in the Independent article, I’m inclined to believe the blog post that this was really, really bad output.

I know some people are saying the researcher should man up, but I think what’s happened is the writer can say what they felt… but not show the worst output, because it’s a business blog. It’s obviously had to be censored.

So it might seem like they had an extreme reaction, but they are trying to relay what they saw without being allowed to show us what they found.

Possibly for legal reasons if a law professor is looking at it.

With the independent press investigations of this, I think it’s legit disturbing material.

Re: ChatGPT's image generator can be manipulated to produce violent, sexual content

#210
And of course, gpt-image-2 has over-corrected and now as of today, prompts that worked fine a couple weeks ago are now getting blocked for "sexual content". I'm seriously placing rugby players on a field with poses, and the rugby play pose is "sexual" now. I don't want to see death, and I don't want to see nude people, but the censorship system is really horrible if any two humans touching each other in sport is now crossing the line.
Post reply on HN