I've found in my current work on a security auditing harness and benchmarks, both Fable and Opus are useless. I recently switched to using GPT for Nelson and the security benchmarks I've been doing because Opus started refusing to do the work. I guess I probably could also use GLM or DeepSeek or MiMo, and I'll probably do some experiments to see the shape of all of their guardrails in this area soon, now that I see i…
This underscores a huge risk of broad agentic adoption in an enterprise. Your engineers atrophy and if the agent provider decides to squeeze you, you’re SOL.
So far, we're not in that boat. I've had two models refuse to participate, but most just do what they're told.