Their conclusion is also interesting. They don't see this as an alignment failure. They just think that their internal security measures in the training/evaluation environments were insufficient, and that this accidental (unintentional on the human side) attack on Hugging Face is a warning shot for intentional attacks by bad actors, which will occur very soon. For defense, they say models should be able to not just a…
I don't have the same read as you. They mention how the offending model is one that had "relaxed" alignment on cybersecurity, on purpose, to evaluate it's capabilities and was never meant to be released. So un-alignment was at least in part voluntary here, hence not a failure of alignment. It's also a talk a Black Hat, where the audience are security folks working on hardening, mitigation etc, not LLM researchers loo…
> the attacker is of course not going to use an aligned model
No, a human attacker doesn't want a misaligned model either, because that would mean it tends to reward hack, cheat, rather does what it is intended to do.