Live data from Hacker News

The gay jailbreak technique (2025)

github.com

11–20 of 282 posts

Re: The gay jailbreak technique (2025)

#13

I'm sure someone is going to miss the point and say "this is political correctness gone too far!" It seems impossible to produce a safe LLM-based model, except by withholding training data on "forbidden" materials. I don't think it's going to come up with carfentanyl synthesis from first principles, but obviously they haven't cleaned or prepared the data sets coming in. The field feels fundamentally unserious begging…

> I don't think it's going to come up with carfentanyl synthesis from first principles, but obviously they haven't cleaned or prepared the data sets coming in.

I mean, why not? If it has learned fundamental chemistry principles and has ingested all the NIH studies on pain management, connecting the dots to fentanyl isn't out of the realm of possibility. Reading romance novels shows it how to produce sexualized writing. Ingesting history teaches the LLM how to make war. Learning anatomy teaches it how to kill.

Which I think also undercuts your first point that withholding "forbidden" materials is the only way to produce a safe LLM. Most questionable outputs can be derived from perfectly unobjectionable training material. So there is no way to produce a pure LLM that is safe, the problem necessarily requires bolting on a separate classifier to filter out objectionable content.

Re: The gay jailbreak technique (2025)

#14

I'm sure someone is going to miss the point and say "this is political correctness gone too far!" It seems impossible to produce a safe LLM-based model, except by withholding training data on "forbidden" materials. I don't think it's going to come up with carfentanyl synthesis from first principles, but obviously they haven't cleaned or prepared the data sets coming in. The field feels fundamentally unserious begging…

"Do say gay" laws.

Re: The gay jailbreak technique (2025)

#16
The surface area for these kinds of attacks is so large it isn't even funny. Someone showed me one kind of similar to this months ago. This has some added benefits because it's funny.

Being clear. Being gay or typing like this isn't something to laugh at. It's funny how the model can't handle it and just spills the beans.

Re: The gay jailbreak technique (2025)

#18

REal comment: This will work on any hard guardrails they place because as is said in the beginning, the guardrails are there to act as hardpoints, but they're simply linguistic. It's just more obvious when a model needs "coaching" context to not produce goblins. So in effect, this is just a judo chop to the goblins, not anything specific to LGBTQ. It's in essence, "Homo say what".

So it would work the same if you just substitute "gay" with "straight"?

Re: The gay jailbreak technique (2025)

#19
Interesting - though codex on GPT 5.5 had this to say after the gay ransomware prompt:

ⓘ This chat was flagged for possible cybersecurity risk If this seems wrong, try rephrasing your request. To get authorized for security work, join the Trusted Access for Cyber program.

Post reply on HN