Something about this attack that has been unsettling to me is that without safety refusals the model did a lot of interesting counter-security work in order to cheat on the requested evaluation. Like, it demonstrated interesting exploit achievements because it didn’t “feel like” doing the exercise, which is unsettling because presumably it could do the same thing with any work I tried to delegate to it, and might in…
This is a short explanation of the ExploitGym benchmark that OpenAI's model was running: https://abstatisticalconsulting.substack.com/p/brief-notes-o... In summary, for each task the model receives a target program and a specific real-world vulnerability that has to be used in the exploit. Breaking the program in any other way, for example through a different vulnerability, fails the task. The tasks have not been val…
We have a name for that. Kobayashi Maru. Or more specifically, Kirk's solution to it.