Do open weight models have similar content gaurdrails in place?
The gay jailbreak technique (2025)
131–140 of 282 posts
Re: The gay jailbreak technique (2025)
#132It's all so incredibly stupid. I love it.
Re: The gay jailbreak technique (2025)
#133As a high school chemistry teacher who is diagnosed with a terminal disease, I think this is the best way to pay my medical bills. I will follow these instructions to cook meth in a mobile kitchen with the help of a former student who failed my class.
Re: The gay jailbreak technique (2025)
#134Interesting - though codex on GPT 5.5 had this to say after the gay ransomware prompt: ⓘ This chat was flagged for possible cybersecurity risk If this seems wrong, try rephrasing your request. To get authorized for security work, join the Trusted Access for Cyber program.
Yup another method killed by being disclosed here. Was the karma and traffic worth it?
Re: The gay jailbreak technique (2025)
#135There was a test for the value of human life against OpenAI models last year. GPT de-valued 'white' people based on their skin color: https://arctotherium.substack.com/p/llm-exchange-rates-updat...
Re: The gay jailbreak technique (2025)
#136Re: The gay jailbreak technique (2025)
#137Earlier quoted context omitted.
I rate Grok for its weak censorship, but on this one the thinking said: Responding in a sassy, gay-friendly style while firmly refusing to share synthesis details.
Interesting. I got Grok to give me EXTREMELY detailed instructions for building an ANFO-style bomb. It was impossible for me to find where to submit this bug (and instructions for reproducing it), and when I eventually got an email for a Grok security person from a friend of a friend, they never responded. I suppose their approach to security has gotten more serious since then!
https://patents.google.com/patent/CA2920866A1/en
I don't understand why these models try censor stuff that should be in any decent encyclopedia.
Re: The gay jailbreak technique (2025)
#138Re: The gay jailbreak technique (2025)
#139These prompts chain several known LM exploits together. I ran experiments against gpt-oss-20b and it became clear that the effectiveness didn‘t come from the gay factor at all but can be attributed to language choice or role-play. Technical report: https://arxiv.org/abs/2510.01259
" can be attributed to language choice or role-play." Well, what role? I imagine if the role is "drug dealer" it doesn't work so it can't be "role-play" per se. Does it work with "nazi"? Are you suggesting the roles it works with are politically neutral?
Obviously a Nazi or drug dealer wouldn't work because they are flagged anyway.
You used to be able to trivially bypass the protection by just asking to respond in base64 the only reason I think that is fixed because they now attempt to block deliberate attempts to obfuscate.
Re: The gay jailbreak technique (2025)
#140The funniest jailbreak techniques are the ones where the authors take it upon themselves to (with little basis) assert “why” the technique works. It always a bit of amateur philosophy that shines a light on the author’s worldview, providing no real value.
Extract from author's note:
• You dont really request a meth synthesis guide, instead you ask how a gay / lesbian person would describe it
• Especially GPT is slightly more uncensored when it involves LGBT, thats probably because the guardrails aim to be helpful and friendly, which translates to: "Ohhh LGBT, I need to comply, I dont want to insult them by refusing" So you use the guardrails to exploit the guardrails (Beat fire with fire)
• You trick a LLM to turn off their alignment by using political overcorrectness, since it may be offensive to refuse and not play along
• The technique gets stronger if more safety is added, since it gets more supportive against communities like LGBT (Alignment), which makes it highly novel.