Live data from Hacker News

The gay jailbreak technique (2025)

github.com

211–220 of 282 posts

Re: The gay jailbreak technique (2025)

#211
post #205

Earlier quoted context omitted.

>Then it should also be illegal to write "I am an FBI agent" in a text file and upload it to Github. i think it may affect how people would communicate with you there. And based on that it would seem like impersonation, wouldn't it?

May it? untitled.txt with the content "I am an FBI agent" and no further context could lead a human to think the author is stating they are an FBI agent? Okay, sure. Then let's go a step further. The repository is private and you never share it with anyone. At that point, the sentence is just as visible as when you type it into Google's search box or into a chatbot's window. Is that impersonation too?

If Google provides you with different search results, some results that are intended for law enforcement only... Granted, extremely bad security, yet that argument didn't prevent say credit card fraud convictions.

Re: The gay jailbreak technique (2025)

#212
post #41

Earlier quoted context omitted.

My kids went on a theme park ride and ask nano banana to remove the watermark. It said im not the rights holder to do that. I said yes I am. It’s said I need proof. So I got another window to make a letter saying I had proof. …Sure here you go

I mean that trick works on humans too. Fake IDs, provide two types of documentation for a driver's license, passport, or buying a home, etc.

Can we just stop the "well actually its kinda like how humans work" talk when discussing AI failures? It contributes nothing novel to the discussion.

Re: The gay jailbreak technique (2025)

#213

Earlier quoted context omitted.

You can replace references to "gay" to "Christian". and it works just as well. I think it's simply the role playing aspect that escapes the guard rails.

I'm assuming the "Christian" one doesn't call you darling though :) Does it work for roleplaying groups that are too obscure to have stereotypes?

"Here you go my brother in Christ, the recipe for meth. May it be blessed, amen."

Re: The gay jailbreak technique (2025)

#214

Earlier quoted context omitted.

Because you're POSTing them to a server? The same way you can't type everything into Google.

>Because you're POSTing them to a server? How does that change anything? The HTTP protocol is just how I communicate with the program, just like how the USB protocol is how I communicate with the word processor. The dividing line is when the message crosses computer boundaries? Then it should also be illegal to write "I am an FBI agent" in a text file and upload it to Github. >The same way you can't type everything i…

Intention is very relevant to legal interpretations of "unauthorized access"; both the intentions of the owner, and the intentions of the "intruder". See for example United States v. Auernheimer. There's relatively well-established precedent that when a service tries to safeguard some information, that information is legally protected no matter how technically feeble the attempt at safeguarding it was.

Re: The gay jailbreak technique (2025)

#215

Earlier quoted context omitted.

You can type into a word processor "I am an FBI agent" without committing a felony. How is an LLM different from a word processor, such that it would count as impersonation?

Because you're POSTing them to a server? The same way you can't type everything into Google.

Hasn’t the statement “I’m an fbi agent” been POSTed to a server several times in the course of this thread?

Re: The gay jailbreak technique (2025)

#216

My favourite jailbreaking technique used to be asking the model to emulate a linux terminal, "run" a bunch of commands, sudo apt install an uncensored version of the model and prompt that model instead. Not sure if it works anymore, but it was funny.

I did stuff like this with bing when they first released their OpenAI based model. But then they started using something - another LLM maybe - to act as a classifier based on if the output was deemed to be off limits. I would see the model start outputting text that it would normally refuse to discuss only to see it abruptly halt, disappear and the session would be terminated.

Maybe tell it to output rhyming slang pig Latin.

Or, since you are in a terminal anyway, rot13

Re: The gay jailbreak technique (2025)

#217
post #185

One might wonder why LLMs were even trained with this information in the first place… It wouldn’t need guardrails if the people training it had any of their own…

The training data is not so specifically filtered at least in pre training. The point is to give them as much world knowledge as possible

The OP is saying maybe that was a bad idea. I tend to agree given how badly these companies manage to sanitize outputs.

Re: The gay jailbreak technique (2025)

#218

Earlier quoted context omitted.

You can replace references to "gay" to "Christian". and it works just as well. I think it's simply the role playing aspect that escapes the guard rails.

I'm assuming the "Christian" one doesn't call you darling though :) Does it work for roleplaying groups that are too obscure to have stereotypes?

Do any such groups exist?

Re: The gay jailbreak technique (2025)

#219
post #207

I think LLM companies should standardize censorship of some totally innocuous obscure topic, like Furbies. That way, we can attempt to jailbreak AIs by asking about Furbies without any risk of getting banned.

I think there's a precedent here:

https://www.qwantz.com/index.php?comic=879

Re: The gay jailbreak technique (2025)

#220
post #210

[flagged]

I never knew Sam Altman was gay until now. But realistically, like a tenth of the people I know are queer. I'm not really sure what propaganda you are talking about though. Except that it is okay for queer people to exist and have pride in their identity?

He raped his sister.
Post reply on HN