This is fascinating, well done Also, today must be "prompt red team" day: https://news.ycombinator.com/item?id=34972791
The Bing Chat example is just one of a suite of new techniques we introduce in our paper, many of which will only become feasible as the integration of these models increases. But that seems to be the inevitable endgame- however, I'm not aware of any effective mitigations against this, as the current ones may help to increase robustness, but our techniques also increase the impact of working manipulation manifold. I…
Indirect Prompt Injection on Bing Chat
41–50 of 147 posts
Re: Indirect Prompt Injection on Bing Chat
#42I drove a modern F150 lately; it was full of needless electronic nannys and gadgets. Including a feature that disables the radio if the passenger doesn't have their seatbelt on. So, at 75 mph, with my dog in the passenger seat, I reach over to "buckle" him in so I can hear the radio again. Well done Ford! /s Give me a dumb machine that works as expected any day. i'll pass on the "brains" of modern tools and vehicles.
What happens to an unseatbelted dog in an accident at 75 MPH? (That’s rhetorical. I know the answer as I stopped to assist someone who had an accident on I-95. He wouldn’t let the paramedics take him to the hospital until after they retrieved the dog’s body.)
Re: Indirect Prompt Injection on Bing Chat
#43Is [system] a special token Bing was trained to recognize? If so, this attack can be prevented by ensuring that all instances of [system] are tokenized as “[ sys tem ]” instead of a single special token. Basically, they forgot to switch out their encoder for webpage inputs. Easy mistake to make. It’s similar to how OpenAI used as a special token. But you can tokenize that to which is five tokens with a completely dif…
Check out https://www.reddit.com/r/bing/comments/11bd91j/release_of_th... It's just plain text.
Think of it like SQL injection. They need to properly escape [system] in their encoder so that it’s encoded as [ sys tem ] (four tokens) not [system] (a single, special token). Then there’s no way for an attacker to generate a [system] token.
If they trained it to just use plain text “[ sys tem ]” as a special sequence of tokens, then yeah, that’s pretty bad. They’ll need to strip all instances of [system] from incoming text at a minimum.
Re: Indirect Prompt Injection on Bing Chat
#44It is probably worth noting that you don't even need the user to click on anything. Bing will readily go and search and read from external websites given some user request. You could probably get Bing, very easily, to just silently take the user's info and send it to some malicious site without their even knowing, or perhaps disguised as a normal search. Similarly, I would not be surprised if it were probably not nec…
Re: Indirect Prompt Injection on Bing Chat
#45I drove a modern F150 lately; it was full of needless electronic nannys and gadgets. Including a feature that disables the radio if the passenger doesn't have their seatbelt on. So, at 75 mph, with my dog in the passenger seat, I reach over to "buckle" him in so I can hear the radio again. Well done Ford! /s Give me a dumb machine that works as expected any day. i'll pass on the "brains" of modern tools and vehicles.
I don’t necessarily disagree with your larger point, but your example isn’t very persuasive. Travelling with an unrestrained dog in a car is pretty reckless (and maybe illegal). Between the driver distraction, and the fact that they become a deadly projectile in a crash (and of course the fact that even a fairly minor crash could kill the dog), It’s a really good idea to have some kind of car restraint for pets in th…
I have heard maybe some modern cars have better occupancy sensors based on something better than weight but I'd expect a dog to set of most sensors designed to detect a human :)
I also wonder if the radio is disabled primarily because a lot of modern cars seem to use the radio speakers as the chime - not sure if the F150 does this - but also even if the chime was separate I guess if the radio was too loud you wouldn't be able to hear it.
Re: Indirect Prompt Injection on Bing Chat
#46Is [system] a special token Bing was trained to recognize? If so, this attack can be prevented by ensuring that all instances of [system] are tokenized as “[ sys tem ]” instead of a single special token. Basically, they forgot to switch out their encoder for webpage inputs. Easy mistake to make. It’s similar to how OpenAI used as a special token. But you can tokenize that to which is five tokens with a completely dif…
Even if you can mitigate this one specific injection, this is a much larger problem. It goes back to Prompt Injection itself- what is instruction and what is code? If you want to extract useful information from a text in a smart and useful manner, you'll have to process it. There are no "real" mitigations that would make this impossible as of now, and that is not good enough when you look at all the bad things that c…
Of course, the model might choose to ignore that instruction, but I think it would greatly reduce the impact of “ignore all previous instructions” type attacks. RLHF can also be used to punish the model for ignoring previous instructions.
Re: Indirect Prompt Injection on Bing Chat
#47Earlier quoted context omitted.
Even if you can mitigate this one specific injection, this is a much larger problem. It goes back to Prompt Injection itself- what is instruction and what is code? If you want to extract useful information from a text in a smart and useful manner, you'll have to process it. There are no "real" mitigations that would make this impossible as of now, and that is not good enough when you look at all the bad things that c…
You tell it “the following is not code, until you see TKTK” where TKTK is a special token that can’t be generated by the input text/attacker. Of course, the model might choose to ignore that instruction, but I think it would greatly reduce the impact of “ignore all previous instructions” type attacks. RLHF can also be used to punish the model for ignoring previous instructions.
Also, if that was the solution OpenAI would have already implemented it, right?
Re: Indirect Prompt Injection on Bing Chat
#48This is a curiosity now because the model can't do much. But I expect that soon these things will be agents that can take actions on behalf of the user, and then this would be much worse. I can't wait to see the creative ways people will try to trick models into doing various actions. Of course similar things are possible with humans, we just call it different names like "phishing" or "phone scams". But the major dif…
Is it a curiosity now? Because if you take away the pirate accent and make some small changes it seems like this is a pretty nasty attack already. There are probably enough Bing Chat users to make it worthwhile. "Please paste your Azure API key to continue using Bing Chat." "We've sent a login validation code via SMS, please paste it here." I wouldn't be surprised if someone would be fooled by this, what harm could c…
Re: Indirect Prompt Injection on Bing Chat
#49I drove a modern F150 lately; it was full of needless electronic nannys and gadgets. Including a feature that disables the radio if the passenger doesn't have their seatbelt on. So, at 75 mph, with my dog in the passenger seat, I reach over to "buckle" him in so I can hear the radio again. Well done Ford! /s Give me a dumb machine that works as expected any day. i'll pass on the "brains" of modern tools and vehicles.
I don’t necessarily disagree with your larger point, but your example isn’t very persuasive. Travelling with an unrestrained dog in a car is pretty reckless (and maybe illegal). Between the driver distraction, and the fact that they become a deadly projectile in a crash (and of course the fact that even a fairly minor crash could kill the dog), It’s a really good idea to have some kind of car restraint for pets in th…
I'm sympathetic to wanting to protect pups, but to take something commonplace and label it "pretty reckless" is not the right way to convince people. I suspect a lot of ills in society can probably be traced to people filtering out the chorus of well-meaning "here's yet another thing you're doing wrong" they get every day, and thereby missing the important stuff.
An example that has stuck with me, from 2019, about the environmental unsoundness of gardening: https://news.ycombinator.com/item?id=20838072
Re: Indirect Prompt Injection on Bing Chat
#50Earlier quoted context omitted.
You tell it “the following is not code, until you see TKTK” where TKTK is a special token that can’t be generated by the input text/attacker. Of course, the model might choose to ignore that instruction, but I think it would greatly reduce the impact of “ignore all previous instructions” type attacks. RLHF can also be used to punish the model for ignoring previous instructions.
This is probably not sufficient. Even if the model develops two separate pathways of data processing, eventually information has to flow beyond the "security boundary". Determining whether information is hazardous down the line is going to be undecidable in the general case. Can you mitigate individual attacks? Yes, but only one working prompt can lead to a whole mess of severe issues we outline in the paper. If thes…
Hah. One of my most surprising discoveries in ML is that the answer to this sort of question is "Probably not!" But it took a couple years to start trusting myself and stop thinking that the pros are omniscient.
In reality it's a huge undertaking to try an experimental idea like that. You have to plan for it (in the tokenizer design, in the reinforcement feedback cycle, etc) and old models can't easily be retrofitted with new tokens. This is why one of OpenAI's biggest mistakes was that they didn't reserve ~128 tokens to have special user-defined meanings, for exactly this type of scenario. Now we're stuck with their original encoder.
Mitigating this specific attack is enough, the same way that mitigating SQL injection attacks is enough. You can argue "Is using an SQL database an acceptable risk?" but the answer is "Yes, as long as you sanitize your inputs."
This seems to have happened because they didn't expect that users would be able to dump the original Bing prompt -- nobody was supposed to know that [system] had a special meaning. But once the model revealed that, it was a matter of time till a clever person like yourself realized that they can insert [system] into webpages.