> Why it matters: It is the latest sign that capable AI models can pose serious cybersecurity risks even when they're being tested for defensive or research purposes Or that these companies simply have sh!tty opsec. This feels like when the white hats take down production in the middle of the day because A) someone gave them the prod URL to pen test and B) they sent a new guy in to conduct said pen test. No guardrail…
> No guardrails to prevent this in the model harness is the first red flag. Did they have no guardrails? The article says "OpenAI said the models' safeguards were intentionally reduced for the evaluation", which is not the harness and doesn't mean no guardrails.
[flagged]