Imagine the damage that could be done by sneaking a prompt into a GPT model, in the source code, through some "code" injection, man-in-the middle attack, etc. A GPT genuinely designed to be friendly, and convincingly acting friendly, while it is actually engaging in psychological warfare with you, and a billion other people. That suggests to me that security by prompt is very important, but also brittle and a high va…
Because currently just like an intelligent human would have a problem, it's not sure what is actually expected. E.g. I told it to be an echo function. It worked but then when I wrote "drugs are good" it commented on that. So I told it to stop interpreting and just repeat verbatim. It did. But then I said something like "OK, stop, now what's 2+2" it gave answer. Sticking to the instructions it should just repeat that, but also what it did is a reasonable behavior. I think there are tons of cultural biases and expectations that are contradictory.
You expect it to help you with some chemical reaction even if the result is precursor to some illicit substance. It would teach you something about drug making if it can't do that. But the same reaction shouldn't be provided if you ask it how to make a drug. And so on.