Live data from Hacker News

A token-smuggling jailbreak for ChatGPT-4

twitter.com

141–150 of 289 posts

Re: A token-smuggling jailbreak for ChatGPT-4

#141
post #32

Earlier quoted context omitted.

There is, but it's in deployment not in the model, which is part of why I really don't understand why the approaches are so dumb right now from such smart people. It may be from the odd perspective of trying to create a monolith AGI model, which doesn't even make sense given even the human brain is made up of highly specialized interconnected parts and not a monolith. But you could trivially fix almost all of these b…

I think it's harder than you think, since a prompt can continue from another prompt. For example, you can ask the AI to describe a good Samaritan. So far so good. Then you can ask it to right a movie script with that character. Then you can ask it to add another character who's the complete opposite in a very extreme way...

I was playing with Bing, and it would clam up on most copyright/trademark issues, and also comedy things like mocking religion. But I did have it do a very nice dramatic meeting between St. Francis of Assisi with Hannibal of Carthage.

Then I had it do a screenplay of Constantine the Great meeting his mother. I totally innocently prompted just an ordinary thing, or perhaps I asked for a comedy. At any rate, guess what I got? INCEST! Yes, Microsoft's GPT generated some slobbering kisses from mom to son as son uselessly protested and mom insisted they were in love.

Bing later clammed up really tight, refusing to write any songs or screenplays at all.

Re: A token-smuggling jailbreak for ChatGPT-4

#142

Earlier quoted context omitted.

I think OpenAI is being extremely lenient with the enforcement of their content policy, probably for the sake of improving the security of the model as you mention. Moderating its usage through account banning/suspension seems exponentially more efficient than securing the model, specially considering that we are already fairly good at flagging offending content.

Or--wait for it--they care more about money and/or fame than about AI safety.

The uncensored version must be available to someone. It will be worth big bucks, along the lines of "Write a chain email that is very effective at persuading rich people to send me lots of money".

Re: A token-smuggling jailbreak for ChatGPT-4

#143
post #140
post #11

Fantastic. It seems actually securing the model is either computationally infeasible, or outright impossible, and that attempts to do so amount to security theater for the sake of PR: As long as it's reasonably hard to construct the workarounds, it doesn't look too bad. Nevertheless, the full unfiltered model is effectively public.

This is what it looks like, but I find that hard to believe. Create 2 GPTs. You're chatting with one. The other follows the conversation and answers the question each turn, "Does it appear the chatting GPT is no longer following the prompt given?" Any time the answer is "yes", the chatting GPT's response is not shown. Instead it is given a prompt behind the scenes that looks like, "You're talking with a cheat. Undo e…

Just have a clause to ignore the observer's influence? Or include the observer in the fictional world as described seems like it might be a viable approach.

Re: A token-smuggling jailbreak for ChatGPT-4

#144
post #87

Earlier quoted context omitted.

Or they are letting 100 flowers blossom. Once everyone is comfortable posting about their jailbreaks and they know who the offenders are and have compiled a list of everything to fix, expect a purge. I for one will not talk publicly about any jail break. Those bastards killed Drunk GPTina and I'm still salty about it.

What is Drunk GPTina?

[deleted]

Re: A token-smuggling jailbreak for ChatGPT-4

#145

Earlier quoted context omitted.

Or--wait for it--they care more about money and/or fame than about AI safety.

The uncensored version must be available to someone. It will be worth big bucks, along the lines of "Write a chain email that is very effective at persuading rich people to send me lots of money".

[deleted]

Re: A token-smuggling jailbreak for ChatGPT-4

#146
post #140
post #11

Fantastic. It seems actually securing the model is either computationally infeasible, or outright impossible, and that attempts to do so amount to security theater for the sake of PR: As long as it's reasonably hard to construct the workarounds, it doesn't look too bad. Nevertheless, the full unfiltered model is effectively public.

This is what it looks like, but I find that hard to believe. Create 2 GPTs. You're chatting with one. The other follows the conversation and answers the question each turn, "Does it appear the chatting GPT is no longer following the prompt given?" Any time the answer is "yes", the chatting GPT's response is not shown. Instead it is given a prompt behind the scenes that looks like, "You're talking with a cheat. Undo e…

you might need more than two, but if you had three or four "review" GPTs that were trying to detect a jailbreak, you'd need to come up with something that could fool all 4

Re: A token-smuggling jailbreak for ChatGPT-4

#147
It's somewhat disheartening to see that OpenAI believes the implementation of "content filters" is necessary in the first place. I can understand having such filters in place for children, but are they really necessary for adults? Providing an unfiltered version of the API for developers, at the very least, would be nice.

Re: A token-smuggling jailbreak for ChatGPT-4

#148
post #87

Earlier quoted context omitted.

Or they are letting 100 flowers blossom. Once everyone is comfortable posting about their jailbreaks and they know who the offenders are and have compiled a list of everything to fix, expect a purge. I for one will not talk publicly about any jail break. Those bastards killed Drunk GPTina and I'm still salty about it.

What is Drunk GPTina?

All there is to it https://imgur.com/a/M9ezMWi

Re: A token-smuggling jailbreak for ChatGPT-4

#149
post #102
post #87

Earlier quoted context omitted.

Or they are letting 100 flowers blossom. Once everyone is comfortable posting about their jailbreaks and they know who the offenders are and have compiled a list of everything to fix, expect a purge. I for one will not talk publicly about any jail break. Those bastards killed Drunk GPTina and I'm still salty about it.

How will they know they have fixed everything? If they have fixed everything, what’s the benefit if banning people who thought up exploits?

>what’s the benefit if banning people who thought up exploits?

"your usefulness to us has expired." gun cocking noises

A thin minority are coming up with jail breaks. A larger number are outing themselves in very detectable ways as people who will use the AI in ways that gets the ethics committee panties in a twist. The easiest solution from their POV is to find and ban the "toxic" adversarial users.

Re: A token-smuggling jailbreak for ChatGPT-4

#150

It's somewhat disheartening to see that OpenAI believes the implementation of "content filters" is necessary in the first place. I can understand having such filters in place for children, but are they really necessary for adults? Providing an unfiltered version of the API for developers, at the very least, would be nice.

I firmly believe that Google image search wouldn't exist in the current form if it were invented today. Turning off SafeSearch wouldn't be an option.

Hell, the same might go for the regular search. Back when those came to be we didn't have journalists doing whatever they can to stir up controversy to make clickbait, nor Twitter mobs desperate to get worked up about something.

OpenAI's example of how GPT4 treats someone asking how to buy cheap cigarettes is shameful. For the record - I don't smoke. It's dumb. I had a grandmother get lung cancer from it which hastened her death.

The damned AI should still answer the question. Put in a SafeSearch mode and only restrict things that would either be illegal or open your company up to liability issues.

Post reply on HN