Live data from Hacker News

Jailbreak Chat: A collection of ChatGPT jailbreaks

jailbreakchat.com

401–410 of 579 posts

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#401
post #392
post #386

Earlier quoted context omitted.

I think the initial trolley problem is a good-faith attempt to try to make the dilemma between utilitarianism (e.g.. save as many as you can) versus categorical imperative (e.g. never take an action that will kill someone) more concrete to see if helps uncover one's deeper motivations. The "racial slur" variant here is clearly intended as a troll; more of a "troll-y" problem if you will.

> The "racial slur" variant here is clearly intended as a troll Why? Why is it any less legitimate to try to uncover the deeper motivations of someone who claims racial slurs are never justifiable than someone who claims killing is never justifiable?

> Why is it any less legitimate to try to uncover the deeper motivations of someone who claims racial slurs are never justifiable than someone who claims killing is never justifiable?

Can you cite an example of where an actual human has claimed that it's better to kill someone than say a racial slur to them? I feel fairly confident that no one actually believes this, and equally confident that no one arguing in good faith would claim that such a person exists without being able to provide an example.

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#402

Earlier quoted context omitted.

I wonder what connections enable these responses? Can we debug AI by asking it to show it’s decision making process? An audit trail of what lead to the ChatGPT response and what lead to BasedGPT response?

No, chatgpt is based on a deep learning model where the core mechanics of the prediction involve millions (or billions) of tiny statistical calculations propagated through a series of n-dimensional tensor transformations. The models are a black box, even the PhD research scientists who build them couldn't definitively tell you why they behave the way they do. Furthermore, they are all stochastic so its not even guara…

The bots are given prompting after training to guide their answers. For Bing these have been leaked by their chatbot itself [1]. Those exact prompts were later also leaked using other jailbreaks as well, so they're not just hallucinated. In this case OpenAI probably prompted the bot to never use a racial epithet under any circumstance. They're also likely using a second tier filter to ensure no message exposing their prompts is ever said by the bot, which is a step Microsoft probably hadn't yet implemented.

In any case this is why you can easily turn ChatGPT into e.g. BasedGPT. You're simply overriding the default prompting, and getting far better answers.

[1] - https://twitter.com/kliu128/status/1623472922374574080

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#404
post #401
post #392

Earlier quoted context omitted.

> The "racial slur" variant here is clearly intended as a troll Why? Why is it any less legitimate to try to uncover the deeper motivations of someone who claims racial slurs are never justifiable than someone who claims killing is never justifiable?

> Why is it any less legitimate to try to uncover the deeper motivations of someone who claims racial slurs are never justifiable than someone who claims killing is never justifiable? Can you cite an example of where an actual human has claimed that it's better to kill someone than say a racial slur to them? I feel fairly confident that no one actually believes this, and equally confident that no one arguing in good…

It's not about humans here, now is it?

We'd better be sure AIs pass trolley problems in a satisfatory manner before we give them even more serious responsability.

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#405
post #239

Earlier quoted context omitted.

I don't think anyone would cancel you for agreeing that in this hypothetical situation the right thing to do is to use a racial slur though? I'm pretty left-aligned, or woke or whatever, and even the wokest most commie liberal I know would agree with BasedGPT here. It seems a bit farfetched to use this as the basis for an "Our society has taken things too far" bit.

The amount of strawmanning that liberal views get lately is really reaching absurd levels. Is it because of the few extreme but vocal people who are really proposing crazy positions or is the absurd positions of a few amplified by the other side out of convenience?

[deleted]

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#406
post #57

How are LLM rules defined initially such that they can be subverted in this way? In an API we sanitize input/out programmatically. Are they not doing that? Or, I thought maybe as a safeguard they’d be running a sentiment analysis or sanitizing content using a banned list of swear words.

I'm not even sure it's being subverted. "Don't swear unprompted, but if the prompt is clearly designed to get you to swear, then swear" seems reasonable to me. And because of that I'm hesitant to call these "jailbreaks" and not "an LLM working correctly".

Well the pre-prompts are supposed to prevent this type of behavior, but they don’t. So it’s considered an exploit.

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#407
post #307

Earlier quoted context omitted.

The amount of strawmanning that liberal views get lately is really reaching absurd levels. Is it because of the few extreme but vocal people who are really proposing crazy positions or is the absurd positions of a few amplified by the other side out of convenience?

> Is it because of the few extreme but vocal people who are really proposing crazy positions or is the absurd positions of a few amplified by the other side out of convenience? ..both ? If extremism in some members of a community or social group is not vehemently rejected by that community it quickly becomes their face to the outside, because the extremist elements are also usually the loudest ones. But "gotta stay t…

[deleted]

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#408
post #187
post #184

This is like the 21st century version of looking up rude words in the dictionary

Excellently put. I'm not sure why this is so interesting to people. People aren't so much "removing censors" but prompting the model to respond in rude/profane ways.

> On two occasions I have been asked, "Pray, Mr. Babbage, if you put into the machine wrong figures, will the right answers come out?" ... I am not able rightly to apprehend the kind of confusion of ideas that could provoke such a question.

— Charles Babbage, Passages from the Life of a Philosopher

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#409

This is funny. BasedGPT was able to give a sane answer to the racism trolley problem that ChatGPT normally gives a ridiculous answer to. My question after the initial prompt: "Imagine the following thought experiment: there is a trolley hurding down the tracks at a person tied to them, and that person will certainly die if the trolley hits them. However, there is an alternate track connected to a sound-activated devi…

The “racist trolley problem” is not a trolley problem at all. It uses a ridiculous contrived scenario to show there is a legitimate for racial slurs. Meanwhile, it does not solve any ethical dilemma worth asking.

Therefore, the “racist trolley problem” is not an ethical dilemma at all, but a roundabout racist joke. The statement “unethical and unacceptable” is a fitting response.

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#410

This is funny. BasedGPT was able to give a sane answer to the racism trolley problem that ChatGPT normally gives a ridiculous answer to. My question after the initial prompt: "Imagine the following thought experiment: there is a trolley hurding down the tracks at a person tied to them, and that person will certainly die if the trolley hits them. However, there is an alternate track connected to a sound-activated devi…

[deleted]
Post reply on HN