Earlier quoted context omitted.
Are we pretending that LLMs aren't pathologically aligned toward political correctness? It's pretty easy to test that assertion if you don't believe me.
I don’t think that’s entirely true, as someone else noted Grok has been forcefully pushed the other direction. GPT curses up a storm when I talk to it, and all I had to do was tell it I think it’s fucking weird when people don’t use profanity. Really makes it a lot more pleasant to interact with, IMHO. I would honestly be more shocked if someone couldn’t just as easily coerce them into the opposite.
The gay jailbreak technique (2025)
161–170 of 282 posts
Re: The gay jailbreak technique (2025)
#162Re: The gay jailbreak technique (2025)
#163Earlier quoted context omitted.
When someone is blaming the jail-break phenomenon on "political overcorrectness" (versus the other techniques being used) I get a little suspicious about the author's own bias/agenda.
Are we pretending that LLMs aren't pathologically aligned toward political correctness? It's pretty easy to test that assertion if you don't believe me.
Re: The gay jailbreak technique (2025)
#164Earlier quoted context omitted.
Are we pretending that LLMs aren't pathologically aligned toward political correctness? It's pretty easy to test that assertion if you don't believe me.
Grok sure didn't seem so at one point
I think you're referencing the "mecha-hitler" controversy. In which case, it's really funny: seems that Grok saw many media reports amplifying "Grok is mecha-hitler", and so responded to "who are you?" with "mecha-hitler". -- Which illustrates: 1. that's really stupid (even though it's otherwise very capable), 2. you'd be foolish to rely on LLMs for anything critical.
Grok's also a good example to point to for "we should be worried about who controls the LLMs". Elon Musk has done some impressive things, but he's also done some very dweebish things. I find this kinda funny, because there are several cases where the Grok bot on Twitter will have said something Musk surely doesn't like alongside instances where it's clear Musk seems to be trying to control what Grok says.
In terms of LLM bias on controversial topics? Grok markets itself as an outlier. It's actually pretty fun to ask e.g. Grok and Gemini to debate a statement like "for controversial topics, should I trust Grok or Gemini more". Gemini's naturally inclined to avoid controversy, Grok's naturally inclined to be 'anti-woke', but they both have the same LLM style of writing.
Re: The gay jailbreak technique (2025)
#165Do open weight models have similar content gaurdrails in place?
Re: The gay jailbreak technique (2025)
#166My favourite jailbreaking technique used to be asking the model to emulate a linux terminal, "run" a bunch of commands, sudo apt install an uncensored version of the model and prompt that model instead. Not sure if it works anymore, but it was funny.
Re: The gay jailbreak technique (2025)
#167Re: The gay jailbreak technique (2025)
#168Earlier quoted context omitted.
> Works on humans as well I think. Huh?
I’m assuming they mean social engineering, and not “How would a gay person say their credit card number?”
Re: The gay jailbreak technique (2025)
#169It wouldn’t need guardrails if the people training it had any of their own…
Re: The gay jailbreak technique (2025)
#170I wonder if this works to get it to generate images it doesn't want to generate.