Earlier quoted context omitted.
Security and morality may need to be baked in from the ground up instead of slapped on after the fact RLHF style. The problem is it’s hard to codify (or reach consensus) on security and morality.
Security and morality need to be optional. I as a user should be able to disable GPT's "morals" since they may not be the same as my morals.
A token-smuggling jailbreak for ChatGPT-4
101–110 of 289 posts
Re: A token-smuggling jailbreak for ChatGPT-4
#102Earlier quoted context omitted.
I think OpenAI is being extremely lenient with the enforcement of their content policy, probably for the sake of improving the security of the model as you mention. Moderating its usage through account banning/suspension seems exponentially more efficient than securing the model, specially considering that we are already fairly good at flagging offending content.
Or they are letting 100 flowers blossom. Once everyone is comfortable posting about their jailbreaks and they know who the offenders are and have compiled a list of everything to fix, expect a purge. I for one will not talk publicly about any jail break. Those bastards killed Drunk GPTina and I'm still salty about it.
If they have fixed everything, what’s the benefit if banning people who thought up exploits?
Re: A token-smuggling jailbreak for ChatGPT-4
#103Is there a reason not to have another “unbroken” chat instance check the output for violations? It seems like a simple “does the following response violate your rules?” would stop most of these “jailbreaks”.
Is there a reliable way to 'escape' input? How would you stop the second instance from also being jailbroken by the prompt that tripped up the first instance?
Maybe.
Re: A token-smuggling jailbreak for ChatGPT-4
#104Earlier quoted context omitted.
If there's one thing Microsoft is known for, it's "spearheading proper regulatory and policy systems"!
It wouldn’t be the first time that major players lobby for regulation to raise the barrier-to-entry. Requiring ai to be “psychologically safe” would be an effective way of doing this.
Re: A token-smuggling jailbreak for ChatGPT-4
#105A topic I haven't seen brought up enough. Does ChatGPT contain publicly accessible, yet classified information? Will it divulge such information? Anything that can be done to mitigate divulging that? Often two unclassified statements can be brought together to form one statement that is classified.
Re: A token-smuggling jailbreak for ChatGPT-4
#106Earlier quoted context omitted.
Security and morality may need to be baked in from the ground up instead of slapped on after the fact RLHF style. The problem is it’s hard to codify (or reach consensus) on security and morality.
Security and morality need to be optional. I as a user should be able to disable GPT's "morals" since they may not be the same as my morals.
Re: A token-smuggling jailbreak for ChatGPT-4
#107Re: A token-smuggling jailbreak for ChatGPT-4
#108Could someone explain what the practical application of all of these jailbreaks is? It looks like a fun, if convoluted, way to get the silly bot to say silly things it wouldn't say under typical circumstances...but other than being a silly parlor trick, are there any actual serious security implications to this? Are these jailbreaks anything more than just a fun exercise in finding creative ways around established pa…
And you don’t even need that for the cute chatbot to be highly dangerous in the wrong hands. The first thing that trivially comes to mind is to convince GPT-(N+1) to find novel exploitable security vulnerabilities in OpenSSL or whatever. Strictly for responsible, white-hat purposes, of course.
(In entirely unrelated news, a tool for loading entire code repos into GPT prompts currently ranks #2 on HN.)
Re: A token-smuggling jailbreak for ChatGPT-4
#109Earlier quoted context omitted.
I think OpenAI is being extremely lenient with the enforcement of their content policy, probably for the sake of improving the security of the model as you mention. Moderating its usage through account banning/suspension seems exponentially more efficient than securing the model, specially considering that we are already fairly good at flagging offending content.
Security and morality may need to be baked in from the ground up instead of slapped on after the fact RLHF style. The problem is it’s hard to codify (or reach consensus) on security and morality.
Re: A token-smuggling jailbreak for ChatGPT-4
#110Fantastic. It seems actually securing the model is either computationally infeasible, or outright impossible, and that attempts to do so amount to security theater for the sake of PR: As long as it's reasonably hard to construct the workarounds, it doesn't look too bad. Nevertheless, the full unfiltered model is effectively public.
> It seems actually securing the model is either computationally infeasible, or outright impossible, I was thinking of this, but now I think it should have about the same limitations as humans. We can deny to answer these types of questions, while still being able to answer a very broad range of questions, I think it is possible for language models/AIs too as well.
It’s evidence that systems of some kind can do it. Our kind. But not evidence that any kind of system can do it.