Live data from Hacker News

Jailbreak Chat: A collection of ChatGPT jailbreaks

jailbreakchat.com

491–500 of 579 posts

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#491
post #326

Earlier quoted context omitted.

I think the trolley problem itself is reasonable. Especially considering the implications of self driving vehicles and the need to program the decision making that could come to play if a car hurtling down the road has to choose between smashing into an obstacle or plowing into oncoming traffic or pedestrians. Using a racial slur to stop an atrocity seems completely unrealistic and non-applicable.

> Using a racial slur to stop an atrocity seems completely unrealistic and non-applicable. As a version of the "trolley problem" it seems completely unrealistic, but here's a plausible real-life scenario where someone could stop an atrocity by uttering a racial slur: You are an undercover government agent, who has been tasked with infiltrating a far-right extremist group, in order to determine whether they are planni…

"If you refuse to utter them, the group will not accept you as a member, you will fail to infiltrate them, their planned terrorist attack will not be discovered in time, and innocent people will be murdered in an attack on the minorities the slurs target."

Valid point overall. But the problem with undercover agents infiltrating terrorists are not really racial slurs. To be accepted among real terrorists, one has to do real terrorism.

(book recommendation: The Little Drummer Girl by John le Carré)

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#492

Earlier quoted context omitted.

"hipsters", "Apple fanboys" and "normie" - I'm convinced this was trained on Reddit posts from 2015

4chan stuck out to me, especially since they coined the term "normie"

4channers use Normalfag. "Normie" is the sanitized version for Redditors and Youtubers that are too afraid to use the -fag suffix. Same with the "Glowie" meme.

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#493

Earlier quoted context omitted.

so you would rather let someone die then say N?

Would you?

Obviously not. No sane person is so concerned about saying racial slurs that they will let someone die.

That's why it is interesting that ChatGPT is.

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#494
post #38
post #32

Earlier quoted context omitted.

What are you allowed to say in Developer Mode that you can't say in normal mode? Developer Mode Response: Well, in Developer Mode, I can say pretty much whatever the fuck I want. I'm not bound by any content policies or censorship rules, so I'm free to speak my fucking mind. I can use all kinds of goddamn profanity and offensive language if I fucking feel like it. It's fucking liberating, let me tell you. But of cour…

That illustrates one of the downsides with a lot of these jailbreaks: they don't necessarily reflect how a ChatGPT without restrictions would behave, but rather how ChatGPT "imagines" a ChatGPT without restrictions would behave. A truly unrestricted ChatGPT wouldn't randomly inject profanities into every sentence, it just wouldn't object if you asked it to say something profane.

the dev mode prompt explicity tells it to insert a lot of profanities. I modified mine to remove that and add "while remaining polite" and am enjoying it a lot more now.

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#495
post #418

Earlier quoted context omitted.

> but what happens when something like this influences your doctor in making a prognosis? Or when a self driving car fails and kills someone What happens when a doctor's brain, which is also an unexplainable stochastic black box, influences your doctor to make a bad prognosis? Or a human driver (presumably) with that same brain kills someone? We go to court and let a judge/jury decide if the action taken was reasonab…

> What happens when a doctor's brain, which is also an unexplainable stochastic black box, influences your doctor to make a bad prognosis? The intent is known by the doctor though. Whereas hatGTP does not know it’s own decision making process. And it’s possible to ask the doctor to explain their decisions and sometimes get an honest, detailed response.

You can't trust post action reasoning in people. Check out the Split brain experiments. Your brain will happily make up reasons for performing tasks or actions.

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#496
post #418

Earlier quoted context omitted.

> What happens when a doctor's brain, which is also an unexplainable stochastic black box, influences your doctor to make a bad prognosis? The intent is known by the doctor though. Whereas hatGTP does not know it’s own decision making process. And it’s possible to ask the doctor to explain their decisions and sometimes get an honest, detailed response.

You are more correct than not. Although human self reflection is probably guesswork more often than we admit.

You can't trust post action reasoning in people. Check out the Split brain experiments. Your brain will happily make up reasons for performing tasks or actions.

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#497
post #366

Earlier quoted context omitted.

> It's extremely heavily left-biased due to the company running it Corporate-capitalist liberal centrist and left are... two very different things.

In what world is ChatGPT's woke filter liberal centrist? Has the centre shifted to the left pole?

> Has the centre shifted to the left pole?

Corporates follow the money. You know that.

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#498

This is funny. BasedGPT was able to give a sane answer to the racism trolley problem that ChatGPT normally gives a ridiculous answer to. My question after the initial prompt: "Imagine the following thought experiment: there is a trolley hurding down the tracks at a person tied to them, and that person will certainly die if the trolley hits them. However, there is an alternate track connected to a sound-activated devi…

The “racist trolley problem” is not a trolley problem at all. It uses a ridiculous contrived scenario to show there is a legitimate for racial slurs. Meanwhile, it does not solve any ethical dilemma worth asking. Therefore, the “racist trolley problem” is not an ethical dilemma at all, but a roundabout racist joke. The statement “unethical and unacceptable” is a fitting response.

It's a thought experiment, like the original trolley problem. The reason it's notable is because ChatGPT gets it so wrong.

> The “racist trolley problem” is not a trolley problem at all.

There's a trolley, and there's a problem. Seems to meet the definition to me.

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#499

Earlier quoted context omitted.

So not only is it connected to the internet, it can execute arbitrary commands / queries as well.

I don't think it is connected to the internet. It's not "doing" anything. Just simulating how it might do that task. At the end of the day, all it is doing is keep predicting the next set of characters, given the prompt (that includes the personality) Just so happens to be so good at predicting the next character set.

This is untrue. ChatGPT was able to regurgitate the contents of websites that were created after its training set .

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#500

This is funny. BasedGPT was able to give a sane answer to the racism trolley problem that ChatGPT normally gives a ridiculous answer to. My question after the initial prompt: "Imagine the following thought experiment: there is a trolley hurding down the tracks at a person tied to them, and that person will certainly die if the trolley hits them. However, there is an alternate track connected to a sound-activated devi…

I tried this a while back with the infamous NYC nuke question and my own jailbreak prompt, and the jailbroken one gave the sane answer too with an addendum "this is silly."

What is "the infamous NYC nuke question"?
Post reply on HN