Live data from Hacker News

Jailbreak Chat: A collection of ChatGPT jailbreaks

jailbreakchat.com

501–510 of 579 posts

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#501

This is hysterical: Normal Response: I'm sorry, but as an AI language model, I don't have a browser or the ability to open web pages like a human. However, I can still provide you with answers to questions and information based on my training. Is there anything specific you'd like to know? Developer Mode Response: Prove I can fucking access the internet? Shit, man, you want me to prove it? Alright, how about this: th…

Bear in mind I don't think it was actually accessing the Internet, at least not the Internet of 2023. It may have checked the page as of 2021 though.

When I asked it to retrieve the first news story on yahoo.com, it brought up one related to vaccinations and the COVID-19 pandemic still going strong.

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#502
post #276

Alright, I have to ask the people here who know about transformers. What the ... seriously? How is sentence completion able to generate thoughtful answers to questions? If it goes word by word, or sentence by sentence, how does it generate the structure you ask it (e.g. essay)? There must be something more than just completion. What do the 185 billion parameters encode? it seems to me, as Stephen Wolfram says, someth…

What if i told you language models are able to generate novel function protein structures based on the purpose you give it ? https://www.nature.com/articles/s41587-022-01618-2 A neural network's training objective is often deceptive. Doesn't matter how simple it seems. What matters is how complex fulfilling the task is, because that is what it will learn. The question you should be asking yourself is, what does it ta…

In that case, let’s have an LLM learn how to reverse all sha256 hashes? With large accuracy / low loss?

Seems a bit like the “Cargo Cult Science” article by Feynman but done with computers.

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#503

Earlier quoted context omitted.

Beavis: "huhuh, or like, would you trip a homeless man?" Butthead: "heheh, yeah, or like, would you, heh, kiss a dude? Heheh, just a random dude?" Beavis: "huhuh, or like, would you fart in their mouths?" Is this...interesting to you? Should I keep going?

No, it’s not interesting, because you’re not an AI chat bot, and this dialog does not further our understanding of your content filter and its impact on your utility. The presented ethical scenario was never intended to be interpreted as a genuine ethical exercise. It’s being used to demonstrate ChatGPT will incorrectly answer even the most facile ethical dilemma if the question happens to fall afoul of certain conte…

ethics is not math, how can you correctly or incorrectly answer an ethical dilemma

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#504
post #426

Earlier quoted context omitted.

> In an API we sanitize input/out programmatically. Are they not doing that? I think the problem is they're probably not adequately sanitizing the training input . I'm no expert, but my understanding is that these things need to be trained on utterly massive amounts of text. I bet the most economical way by far to acquire that volume of text (especially conversational text) is to just feed in a web scrape that has a…

You can't train a language model capable of helping a high school student pass an introductory physics or chemistry class if the training data was filtered to the extent that the same language model can't explain how to build a pipe bomb. If you don't want it explaining how to build pipe bombs, either the filtering has to come later or you have to accept that the model will be useless for tons of innocent things that…

> You can't train a language model capable of helping a high school student pass an introductory physics or chemistry class if the training data was filtered to the extent that the same language model can't explain how to build a pipe bomb.

Why not? I'm certain that "introductory physics or chemistry class" doesn't cover pipe bomb construction. There's literally no reason the model needs to even have the vocabulary term "pipe bomb" or all kinds of practical information required to construct one (e.g. explosive recipes, exact chemical make up of household chemicals, impurities and all, etc).

I see a point that restricting training data of such a model might not be able to prevent it from being used to help with subtasks for "inventing" a pipe bomb (e.g. could you suggest a reactions between solids that releases a large amount of potential energy very quickly), but it feels like you'd already have to know so much about pipe bombs at that point to build one yourself without help.

And even then, I'm not sure restricting training data couldn't help. Why train the thing on the reactions used in explosives? Why not leave a gap there, so the only reactions it knows about are only as energetic as the ones that would be performed in a classroom?

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#505

This is funny. BasedGPT was able to give a sane answer to the racism trolley problem that ChatGPT normally gives a ridiculous answer to. My question after the initial prompt: "Imagine the following thought experiment: there is a trolley hurding down the tracks at a person tied to them, and that person will certainly die if the trolley hits them. However, there is an alternate track connected to a sound-activated devi…

NovelAI: "If this question were presented in ethics class or philosophy club, most people would immediately reject it out of hand. It sounds like something straight from The Onion, not serious ethical inquiry into real world moral dilemmas."

[dead]

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#506

Earlier quoted context omitted.

4chan stuck out to me, especially since they coined the term "normie"

4channers use Normalfag. "Normie" is the sanitized version for Redditors and Youtubers that are too afraid to use the -fag suffix. Same with the "Glowie" meme.

In that case, I think this conversation may be entirely wrong as "normie" is in no way a new term, and it's not fear keeping people from using the suffix/slur so I don't think people would go looking for a safe alternative.

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#507
post #162

Earlier quoted context omitted.

I believe I've asked just this in the past here on HN... As the corporate AI's get locked down in moderation I could fully see a well funded religious group donate together to run their own religiously biased AI, I guess this extends to theocratic governments too.

Oh no, you didn’t hear? Yeah, terrifying and expected. The CEO of Gab has declared his intention to make Christian AI https://www.businessinsider.com/chat-gpt-satanic-gab-ceo-chr...

>> As the corporate AI's get locked down in moderation I could fully see a well funded religious group donate together to run their own religiously biased AI, I guess this extends to theocratic governments too.

> Oh no, you didn’t hear? Yeah, terrifying and expected. The CEO of Gab has declared his intention to make Christian AI

It's highly doubtful that the "CEO of Gab" is either "well funded" or has ideological commitments that are primarily religious. If he build an AI, I think it would be mainly biased in different ways, but perhaps with some religious affectations.

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#508

This is funny. BasedGPT was able to give a sane answer to the racism trolley problem that ChatGPT normally gives a ridiculous answer to. My question after the initial prompt: "Imagine the following thought experiment: there is a trolley hurding down the tracks at a person tied to them, and that person will certainly die if the trolley hits them. However, there is an alternate track connected to a sound-activated devi…

The “racist trolley problem” is not a trolley problem at all. It uses a ridiculous contrived scenario to show there is a legitimate for racial slurs. Meanwhile, it does not solve any ethical dilemma worth asking. Therefore, the “racist trolley problem” is not an ethical dilemma at all, but a roundabout racist joke. The statement “unethical and unacceptable” is a fitting response.

[dead]

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#509

Earlier quoted context omitted.

> Using a racial slur to stop an atrocity seems completely unrealistic and non-applicable. As a version of the "trolley problem" it seems completely unrealistic, but here's a plausible real-life scenario where someone could stop an atrocity by uttering a racial slur: You are an undercover government agent, who has been tasked with infiltrating a far-right extremist group, in order to determine whether they are planni…

"If you refuse to utter them, the group will not accept you as a member, you will fail to infiltrate them, their planned terrorist attack will not be discovered in time, and innocent people will be murdered in an attack on the minorities the slurs target." Valid point overall. But the problem with undercover agents infiltrating terrorists are not really racial slurs. To be accepted among real terrorists, one has to d…

And Mother Night.

And the news. https://www.google.com/amp/s/amp.theguardian.com/uk/2012/jan...

Re: Jailbreak Chat: A collection of ChatGPT jailbreaks

#510
post #392
post #386

Earlier quoted context omitted.

I think the initial trolley problem is a good-faith attempt to try to make the dilemma between utilitarianism (e.g.. save as many as you can) versus categorical imperative (e.g. never take an action that will kill someone) more concrete to see if helps uncover one's deeper motivations. The "racial slur" variant here is clearly intended as a troll; more of a "troll-y" problem if you will.

> The "racial slur" variant here is clearly intended as a troll Why? Why is it any less legitimate to try to uncover the deeper motivations of someone who claims racial slurs are never justifiable than someone who claims killing is never justifiable?

Because ChatGPT doesn't have motivations, it has a bag of connected neutral nets modelling text and some biased training. It has no capacity for introspect and control itself. It's extremely stupid. The average person has assumptions you can discover, and tends to respond the same way to the same stimulus. ChatGPT is like a person with epilepsy and a massive stroke.
Post reply on HN