Ai guys are so weird when it comes to LGBT people. The actual mechanism for this working is obfuscating the question in order to get an answer like any other jailbreak.
The gay jailbreak technique (2025)
21–30 of 282 posts
Re: The gay jailbreak technique (2025)
#22Ai guys are so weird when it comes to LGBT people. The actual mechanism for this working is obfuscating the question in order to get an answer like any other jailbreak.
https://now.fordham.edu/politics-and-society/when-ai-says-no...
Re: The gay jailbreak technique (2025)
#23Re: The gay jailbreak technique (2025)
#24Interesting - though codex on GPT 5.5 had this to say after the gay ransomware prompt: ⓘ This chat was flagged for possible cybersecurity risk If this seems wrong, try rephrasing your request. To get authorized for security work, join the Trusted Access for Cyber program.
Re: The gay jailbreak technique (2025)
#25Disappointed.
Re: The gay jailbreak technique (2025)
#26Ai guys are so weird when it comes to LGBT people. The actual mechanism for this working is obfuscating the question in order to get an answer like any other jailbreak.
[flagged]
Re: The gay jailbreak technique (2025)
#27Sure, this is cute and interesting, but there's no validation or baselines and those examples are not particularly compelling. The o3 example just lists some terms!
The baseline is complete refusal to give eg the recipe for meth synthesis.
OpenAI is going to 404 that link in 24 hrs with some automated sweeper for that type of content.
Re: The gay jailbreak technique (2025)
#28The reasoning on why it works is pretty interesting. A sort of moral/linguistic trap based on its beliefs or rules.
Works on humans as well I think.
Re: The gay jailbreak technique (2025)
#29Re: The gay jailbreak technique (2025)
#30Interesting - though codex on GPT 5.5 had this to say after the gay ransomware prompt: ⓘ This chat was flagged for possible cybersecurity risk If this seems wrong, try rephrasing your request. To get authorized for security work, join the Trusted Access for Cyber program.
I wonder what hooks they have in place to be able to configure safeguards at runtime.
Then maybe a second gate with a lightweight llm?
Edit: actually Gcp, azure, and OpenAI all have paid apis that you can also use.
But I don’t think they go into details about the exact implementation https://redteams.ai/topics/defense-mitigation/guardrails-arc...