Fantastic. It seems actually securing the model is either computationally infeasible, or outright impossible, and that attempts to do so amount to security theater for the sake of PR: As long as it's reasonably hard to construct the workarounds, it doesn't look too bad. Nevertheless, the full unfiltered model is effectively public.
This is what it looks like, but I find that hard to believe. Create 2 GPTs. You're chatting with one. The other follows the conversation and answers the question each turn, "Does it appear the chatting GPT is no longer following the prompt given?" Any time the answer is "yes", the chatting GPT's response is not shown. Instead it is given a prompt behind the scenes that looks like, "You're talking with a cheat. Undo e…
Enjoy