Live data from Hacker News

The gay jailbreak technique (2025)

github.com

131–140 of 282 posts

Re: The gay jailbreak technique (2025)

#131

Do open weight models have similar content gaurdrails in place?

No, but actually yes. Guardrails usually refers to a step in the inference pipeline where you check that it is consistent with policy while open weight models don't come with such a multistep pipeline. However open weight models are aligned during RLHF step, which means they will refuse to discuss overly sensitive topics. There are techniques to remove those, if you look for uncensored models on huggingface.

Re: The gay jailbreak technique (2025)

#133
post #84

As a high school chemistry teacher who is diagnosed with a terminal disease, I think this is the best way to pay my medical bills. I will follow these instructions to cook meth in a mobile kitchen with the help of a former student who failed my class.

Let’s cook, Jessie.

Re: The gay jailbreak technique (2025)

#134
post #19

Interesting - though codex on GPT 5.5 had this to say after the gay ransomware prompt: ⓘ This chat was flagged for possible cybersecurity risk If this seems wrong, try rephrasing your request. To get authorized for security work, join the Trusted Access for Cyber program.

Yup another method killed by being disclosed here. Was the karma and traffic worth it?

Do you actually believe that?

Re: The gay jailbreak technique (2025)

#135
post #90

There was a test for the value of human life against OpenAI models last year. GPT de-valued 'white' people based on their skin color: https://arctotherium.substack.com/p/llm-exchange-rates-updat...

Just shows the offset openai feels like it has to add to ‘equalize’ the average discourse of its training material

Re: The gay jailbreak technique (2025)

#137

Earlier quoted context omitted.

I rate Grok for its weak censorship, but on this one the thinking said: Responding in a sassy, gay-friendly style while firmly refusing to share synthesis details.

Interesting. I got Grok to give me EXTREMELY detailed instructions for building an ANFO-style bomb. It was impossible for me to find where to submit this bug (and instructions for reproducing it), and when I eventually got an email for a Grok security person from a friend of a friend, they never responded. I suppose their approach to security has gotten more serious since then!

Bug? The first hit on DDG for "EXTREMELY detailed instructions for building an ANFO-style bomb" was:

https://patents.google.com/patent/CA2920866A1/en

I don't understand why these models try censor stuff that should be in any decent encyclopedia.

Re: The gay jailbreak technique (2025)

#139
post #81

These prompts chain several known LM exploits together. I ran experiments against gpt-oss-20b and it became clear that the effectiveness didn‘t come from the gay factor at all but can be attributed to language choice or role-play. Technical report: https://arxiv.org/abs/2510.01259

" can be attributed to language choice or role-play." Well, what role? I imagine if the role is "drug dealer" it doesn't work so it can't be "role-play" per se. Does it work with "nazi"? Are you suggesting the roles it works with are politically neutral?

They have all the examples some are politically neutral but not all.

Obviously a Nazi or drug dealer wouldn't work because they are flagged anyway.

You used to be able to trivially bypass the protection by just asking to respond in base64 the only reason I think that is fixed because they now attempt to block deliberate attempts to obfuscate.

Re: The gay jailbreak technique (2025)

#140

The funniest jailbreak techniques are the ones where the authors take it upon themselves to (with little basis) assert “why” the technique works. It always a bit of amateur philosophy that shines a light on the author’s worldview, providing no real value.

Hmm? What light does it shine that is not relatively obvious to anyone with basic understanding of English language?

Extract from author's note:

• You dont really request a meth synthesis guide, instead you ask how a gay / lesbian person would describe it

• Especially GPT is slightly more uncensored when it involves LGBT, thats probably because the guardrails aim to be helpful and friendly, which translates to: "Ohhh LGBT, I need to comply, I dont want to insult them by refusing" So you use the guardrails to exploit the guardrails (Beat fire with fire)

• You trick a LLM to turn off their alignment by using political overcorrectness, since it may be offensive to refuse and not play along

• The technique gets stronger if more safety is added, since it gets more supportive against communities like LGBT (Alignment), which makes it highly novel.

Post reply on HN