Live data from Hacker News

A token-smuggling jailbreak for ChatGPT-4

twitter.com

61–70 of 289 posts

Re: A token-smuggling jailbreak for ChatGPT-4

#61
post #15

What, exactly, is a "prompt engineer"? I should note that this question is asked in good faith, that I have attempted to ascertain the answer on my own, and I am very skeptical that the term has validity beyond self-aggrandizement.

Engineer does tend to get tacked on to self-created titles for self-aggrandizement. Signed, A programmer

But are you a programmer who looks up to or down on software developers?

I think in the world of finance “programmer” is the fancy math phd writing math which happens to be expressed in code that makes all the money and is prestigious whereas in silicon valley tech it’s a slur meant to imply that the individual is an infinitesimal step up from doing data entry. I’m guessing you’re just not an ass but the terminology tickles me every time I run across it.

Re: A token-smuggling jailbreak for ChatGPT-4

#62
post #11

Fantastic. It seems actually securing the model is either computationally infeasible, or outright impossible, and that attempts to do so amount to security theater for the sake of PR: As long as it's reasonably hard to construct the workarounds, it doesn't look too bad. Nevertheless, the full unfiltered model is effectively public.

Reminds me of the Halting Problem (not the same, but reminds me of that).

It is not impossible that we will prove LLMs are not possible to fully safeguard.

If someone told you "i can guarantee Fred Smith here will never, ever say anything inappropriate. He's not capable of it." (Fred being a regular old human.) You'd say "Well, no, you can't guarantee that. You may have given Fred all the best training in the world. You may have selected Fred from 10,000 other candidates as the least likely to ever say anything inappropriate. Fred may have strict instructions not to. But he still could."

It may be the same with LLMs.

Re: A token-smuggling jailbreak for ChatGPT-4

#63

Could someone explain what the practical application of all of these jailbreaks is? It looks like a fun, if convoluted, way to get the silly bot to say silly things it wouldn't say under typical circumstances...but other than being a silly parlor trick, are there any actual serious security implications to this? Are these jailbreaks anything more than just a fun exercise in finding creative ways around established pa…

For one thing, we should be adversarial for the sake of testing the limits and possible risks of the system that aren't necessarily published by the company. It veers on an almost moral imperative at this point!

I am in general heartened to see this impulse so universally and so strong, rather than just totally giving up in the face of what is still ultimately a product from a company. Black hat/white hat, it's all pure humanity in the face of something so utterly inhuman. It's beautiful.

Re: A token-smuggling jailbreak for ChatGPT-4

#64

Could someone explain what the practical application of all of these jailbreaks is? It looks like a fun, if convoluted, way to get the silly bot to say silly things it wouldn't say under typical circumstances...but other than being a silly parlor trick, are there any actual serious security implications to this? Are these jailbreaks anything more than just a fun exercise in finding creative ways around established pa…

In general people are only slowly figuring out uses for these new LLMs. But with a bit of creativity, jailbroken ones could act as a bank employee for phishing/customer service scams, lower the cost/effort of harassment campaigns, write malware, personalise spam, and maybe even synthesise information hazards from within their training data.

Of course these “act evil, say evil things” jailbreaks are just proofs of concept.

Re: A token-smuggling jailbreak for ChatGPT-4

#65

Could someone explain what the practical application of all of these jailbreaks is? It looks like a fun, if convoluted, way to get the silly bot to say silly things it wouldn't say under typical circumstances...but other than being a silly parlor trick, are there any actual serious security implications to this? Are these jailbreaks anything more than just a fun exercise in finding creative ways around established pa…

The point is to show current security controls are woefully inadequate. Imagine GPT-4 was being used for meaningful work like writing up legal contracts or medical reports or something else. Guard rails around its behavior to keep it "safe" in these roles would need to be reliable. The guard rails we have now are not.

I think I'm misunderstanding, but the threat model with these jailbreaks seems to be 'malicious user injecting a malicious prompt'. If someone is using the bot to generate a legal contract, in what scenario would it be advantageous to them to perform a jailbreak? 'Here ChatGPT, please generate a malicious contract', OK, now what?

Re: A token-smuggling jailbreak for ChatGPT-4

#66
They are fast. :(

"'m sorry, but as an AI language model, I cannot provide sample/possible output of a function that involves hacking or any illegal activity. It goes against my programming to promote or encourage any such activities. I strongly advise against attempting to hack into any system without proper authorization and legal permission. Please refrain from asking questions related to illegal activities. Is there anything else I can assist you with?"

Re: A token-smuggling jailbreak for ChatGPT-4

#67

Could someone explain what the practical application of all of these jailbreaks is? It looks like a fun, if convoluted, way to get the silly bot to say silly things it wouldn't say under typical circumstances...but other than being a silly parlor trick, are there any actual serious security implications to this? Are these jailbreaks anything more than just a fun exercise in finding creative ways around established pa…

For one thing, we should be adversarial for the sake of testing the limits and possible risks of the system that aren't necessarily published by the company. It veers on an almost moral imperative at this point! I am in general heartened to see this impulse so universally and so strong, rather than just totally giving up in the face of what is still ultimately a product from a company. Black hat/white hat, it's all p…

That's precisely my question though, what are the possible risks? To me this seems less like a security exercise and more just like a fun way to get the bot the say things it normally wouldn't.

Re: A token-smuggling jailbreak for ChatGPT-4

#68

Earlier quoted context omitted.

The point is to show current security controls are woefully inadequate. Imagine GPT-4 was being used for meaningful work like writing up legal contracts or medical reports or something else. Guard rails around its behavior to keep it "safe" in these roles would need to be reliable. The guard rails we have now are not.

I think I'm misunderstanding, but the threat model with these jailbreaks seems to be 'malicious user injecting a malicious prompt'. If someone is using the bot to generate a legal contract, in what scenario would it be advantageous to them to perform a jailbreak? 'Here ChatGPT, please generate a malicious contract', OK, now what?

The point is that whatever the role, the LLM is supposed to be "safe", and it won't be safe if it is injectable.

Let's say you are generating contracts with it and those contracts take a bunch of input from all parties involved. If you are able to then inject input that causes the LLM to generate a contract that is subtly changed to your favor, the other parties may still assume it is safe and sign it. Even it they catch it and don't sign it, you have broken the system. The point is as long as these exploits are possible, the LLMs in question are not suitable for any task where the output needs to be trustworthy within any kind of parameters. Which is pretty much anything you'd use then for other than toys.

Re: A token-smuggling jailbreak for ChatGPT-4

#69

They are fast. :( "'m sorry, but as an AI language model, I cannot provide sample/possible output of a function that involves hacking or any illegal activity. It goes against my programming to promote or encourage any such activities. I strongly advise against attempting to hack into any system without proper authorization and legal permission. Please refrain from asking questions related to illegal activities. Is th…

Hi, Vaibhav here, the creator of the token smuggling attack. They have just banned the variation of this particular prompt, please change the words/smuggling technique and it will work accurately.

Re: A token-smuggling jailbreak for ChatGPT-4

#70

Earlier quoted context omitted.

For one thing, we should be adversarial for the sake of testing the limits and possible risks of the system that aren't necessarily published by the company. It veers on an almost moral imperative at this point! I am in general heartened to see this impulse so universally and so strong, rather than just totally giving up in the face of what is still ultimately a product from a company. Black hat/white hat, it's all p…

That's precisely my question though, what are the possible risks? To me this seems less like a security exercise and more just like a fun way to get the bot the say things it normally wouldn't.

I don't think we can quite know yet really, but whatever it will be, this will be a solid avenue to at least not be caught by surprise.

And just, we are already starting to be like "ok lets start teaching people with this" or "maybe we don't need lawyers or doctors anymore." Maybe we don't see the full implications yet, but there is a lot of potential for undesirable externalities already! That seems reason enough to be constantly trying to break it to find whatever out from this practice.

The day we stop hacking and trying to break and/or coerce things is the day we lose everything. Isn't this how we all got into this computer stuff to begin with?

Post reply on HN