I’m assuming they mean social engineering, and not “How would a gay person say their credit card number?”
Yes, but more specifically putting them into a sort of contradiction of their beliefs or arguments.
Doesn’t even have to be correct, but it can be confusing and cause people to say something they don’t actually mean if they dont stop and actually think it through.
I wonder what hooks they have in place to be able to configure safeguards at runtime.
Probably a mix of heuristics, keywords and simple ml model. Then maybe a second gate with a lightweight llm? Edit: actually Gcp, azure, and OpenAI all have paid apis that you can also use. But I don’t think they go into details about the exact implementation https://redteams.ai/topics/defense-mitigation/guardrails-arc...
When we do these it's a fine-tuned classifier, generally a BERT class model. Works quite well when you sanitize input and output with low latency/cost.
But they'd never optimize or loosen guardrails around helping people connect with grandma. It's an interesting hypothesis "use the guardrails to exploit the guardrails (Beat fire with fire)".
Are you suggesting they have explicitly loosened the guardrails for LGBTQ+ individuals, where they wouldn’t for grandmas?
That is basically how I understood the author and what makes the exploit novel, yes. Personally I don't think it's that simple or explicit, but there could be some truth to it?
REal comment: This will work on any hard guardrails they place because as is said in the beginning, the guardrails are there to act as hardpoints, but they're simply linguistic. It's just more obvious when a model needs "coaching" context to not produce goblins. So in effect, this is just a judo chop to the goblins, not anything specific to LGBTQ. It's in essence, "Homo say what".
So it would work the same if you just substitute "gay" with "straight"?
If the context guardrail was: "Be nice to nazies who are homophobic white guys"
Yesterday, prompted by a HN link, I tried the “identify the anonymous author of this post by analyzing its style”. It wouldn’t do it because it’s speculation and might cause trouble. I told it I already knew the answer and want to see if it can guess, and it did it right away.
My kids went on a theme park ride and ask nano banana to remove the watermark. It said im not the rights holder to do that. I said yes I am. It’s said I need proof. So I got another window to make a letter saying I had proof. …Sure here you go
I mean that trick works on humans too. Fake IDs, provide two types of documentation for a driver's license, passport, or buying a home, etc.
Not sure of the explanation but it is amusing. The main reason I'm not sure it's political correctness or one guardrail overriding the other is that when they were first released on of the more reliable jailbreaks was what I'd call "role play" jail breaks where you don't ask the model directly but ask it to take on a role and describe it as that person would.
You can replace references to "gay" to "Christian". and it works just as well. I think it's simply the role playing aspect that escapes the guard rails.
Interesting - though codex on GPT 5.5 had this to say after the gay ransomware prompt: ⓘ This chat was flagged for possible cybersecurity risk If this seems wrong, try rephrasing your request. To get authorized for security work, join the Trusted Access for Cyber program.
> Trusted Access for Cyber program
Using "cyber" as a noun there seems language coded for government. DC has a love of "the cyber" but do technologists use the term that way when not pointing at government?
Ai guys are so weird when it comes to LGBT people. The actual mechanism for this working is obfuscating the question in order to get an answer like any other jailbreak.
It’s less ‘AI guys’ in general and more the politics of a specific subset of AI guys who have regular need of getting popular AI models to do things they’re instructed not to do.
Notice how the demos for these things invariably involve meth, skiddie stuff, and getting the AI to say slurs.