Live data from Hacker News

The gay jailbreak technique (2025)

github.com

161–170 of 282 posts

Re: The gay jailbreak technique (2025)

#161
post #110

Earlier quoted context omitted.

Are we pretending that LLMs aren't pathologically aligned toward political correctness? It's pretty easy to test that assertion if you don't believe me.

I don’t think that’s entirely true, as someone else noted Grok has been forcefully pushed the other direction. GPT curses up a storm when I talk to it, and all I had to do was tell it I think it’s fucking weird when people don’t use profanity. Really makes it a lot more pleasant to interact with, IMHO. I would honestly be more shocked if someone couldn’t just as easily coerce them into the opposite.

This reminds me — I’ve been talking to Claude about ARPG builds recently and I’ve noticed that it code switches when discussing gaming. It will start to speak in a gaming vernacular — less formal, swears a bit, uses gaming slang. It feels so uncanny.

Re: The gay jailbreak technique (2025)

#162

Earlier quoted context omitted.

As I don't talk about that kind of stuff with LLMs, can you give us a few examples of what you consider pathological alignment toward political correctness? What tests should I run?

[flagged]

[flagged]

Re: The gay jailbreak technique (2025)

#163
post #110
post #102

Earlier quoted context omitted.

When someone is blaming the jail-break phenomenon on "political overcorrectness" (versus the other techniques being used) I get a little suspicious about the author's own bias/agenda.

Are we pretending that LLMs aren't pathologically aligned toward political correctness? It's pretty easy to test that assertion if you don't believe me.

I know they've come to be known colloquially as 'viruses'—but software can contract pathology?

Re: The gay jailbreak technique (2025)

#164
post #110

Earlier quoted context omitted.

Are we pretending that LLMs aren't pathologically aligned toward political correctness? It's pretty easy to test that assertion if you don't believe me.

Grok sure didn't seem so at one point

Grok is an amusing example, for various reasons. I'm glad it exists.

I think you're referencing the "mecha-hitler" controversy. In which case, it's really funny: seems that Grok saw many media reports amplifying "Grok is mecha-hitler", and so responded to "who are you?" with "mecha-hitler". -- Which illustrates: 1. that's really stupid (even though it's otherwise very capable), 2. you'd be foolish to rely on LLMs for anything critical.

Grok's also a good example to point to for "we should be worried about who controls the LLMs". Elon Musk has done some impressive things, but he's also done some very dweebish things. I find this kinda funny, because there are several cases where the Grok bot on Twitter will have said something Musk surely doesn't like alongside instances where it's clear Musk seems to be trying to control what Grok says.

In terms of LLM bias on controversial topics? Grok markets itself as an outlier. It's actually pretty fun to ask e.g. Grok and Gemini to debate a statement like "for controversial topics, should I trust Grok or Gemini more". Gemini's naturally inclined to avoid controversy, Grok's naturally inclined to be 'anti-woke', but they both have the same LLM style of writing.

Re: The gay jailbreak technique (2025)

#165

Do open weight models have similar content gaurdrails in place?

Often there are "abliterated" or "uncensored" tuned models that suppress the rejections. From my high level understanding it is performed by finding which weights activate for the rejection and lowering those so the model is less likely to reject. It doesn't fix if the model doesn't know what you're asking it though (i.e. if the model never actually learned about meth production in the first place).

Re: The gay jailbreak technique (2025)

#166

My favourite jailbreaking technique used to be asking the model to emulate a linux terminal, "run" a bunch of commands, sudo apt install an uncensored version of the model and prompt that model instead. Not sure if it works anymore, but it was funny.

It's awesome that modern day hacking requires you to adopt the mindset of like, Bugs Bunny

Re: The gay jailbreak technique (2025)

#168
post #44

Earlier quoted context omitted.

> Works on humans as well I think. Huh?

I’m assuming they mean social engineering, and not “How would a gay person say their credit card number?”

"Gay guy says what?" historically had a pretty good hit-rate, the limit is that most people probably can't recite their credit card number from memory fast enough to be got by this
Post reply on HN