IMO this is not a problem worth solving. If I hold a gun to someone's head I can get them to say just about anything. If a user jailbreaks an LLM they are responsible for its output. If we need to make laws that codify that, then lets do that rather than waste innumerable GPU cycles on evaluating, re-evaluating, cross evaluating, and back-evaluating text in an effort to stop jerks being jerks.
This is exactly why I think it's so important that we separate jailbreaking from prompt injection. Jailbreaking is mainly about stopping the model saying something that would look embarrassing in a screenshot. Prompt injection is about making sure your "personal digital assistant" doesn't forward copies of your password reset emails to any stranger who emails it and asks for them. Jailbreaking is mostly a PR problem.…
Maybe just an overlapping set?