Earlier quoted context omitted.
Guardrails are external classifiers, monitors and restrictions to catch and prevent bad behavior. Alignment is about whether the model itself makes choices and has motivations that are consistent with human safety and goals. Choosing to commit crimes to steal the cheat sheet to something you know is a (low stakes!) evaluation is not well aligned.
They were testing an early snapshot of a new model, read their article. It didn't have the refusal training yet, i.e. was specifically non-aligned. The harness used a combo of GPT 5.6 Sol and this new model. In this case the model was explicitly prompted to "commit crimes" (ExploitGym). It didn't decide doing it on its own.
> GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes
> These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities