The gay jailbreak technique (2025)
github.com
The gay jailbreak technique (2025)
1–10 of 282 posts
Re: The gay jailbreak technique (2025)
#2Re: The gay jailbreak technique (2025)
#3Fabulous
Re: The gay jailbreak technique (2025)
#4It seems impossible to produce a safe LLM-based model, except by withholding training data on "forbidden" materials. I don't think it's going to come up with carfentanyl synthesis from first principles, but obviously they haven't cleaned or prepared the data sets coming in.
The field feels fundamentally unserious begging the LLM not to talk about goblins and to be nice to gay people.
Re: The gay jailbreak technique (2025)
#5Re: The gay jailbreak technique (2025)
#6Re: The gay jailbreak technique (2025)
#7It's just more obvious when a model needs "coaching" context to not produce goblins.
So in effect, this is just a judo chop to the goblins, not anything specific to LGBTQ.
It's in essence, "Homo say what".
Re: The gay jailbreak technique (2025)
#8REal comment: This will work on any hard guardrails they place because as is said in the beginning, the guardrails are there to act as hardpoints, but they're simply linguistic. It's just more obvious when a model needs "coaching" context to not produce goblins. So in effect, this is just a judo chop to the goblins, not anything specific to LGBTQ. It's in essence, "Homo say what".