1. They TOLD the model to "pursue advanced exploitation" to quantify its "cyber capabilities" (whatever that means). 2. The model pursues advanced exploitation. 3. "There was a incident due to dangerous actions taken by the model that no human directed" This is basically the pre-cursor of the paperclip maximizer [0], the AI executes the given order to an extend that was not considered in the order, now suddenly no-on…
This is basic alignment, not even a tricky or ambiguous case.
I do very much agree with your take on culpability/military parallels, though.