Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]
cdn.prod.website-files.com
Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]
1–10 of 65 posts
Re: Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]
#2Even that guarantees almost nothing about real alignment (making the AIs want to predict and behave how we would have wanted them to behave).
Re: Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]
#3Re: Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]
#4Re: Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]
#5I am somewhat confident that right now we have crossed a threshold of model capability that we will continue to see such breaches and unsanctioned actions by models in the coming months, some of which would be out in the wild, until someone comes up with some really robust control (keeping the AIs on leash) technique that adequately enforces the sanctioned actions. Even that guarantees almost nothing about real align…
Re: Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]
#6Re: Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]
#7Re: Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]
#8“As a result, the AI agent created a GitHub account…” Why do we have captchas again?
Re: Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]
#9The developer safeguards were off, the models had unfettered access to the internet, and were solving cybersecurity challenges. This happened _after_ the recent OpenAI incident, and the subsequent Anthropic one. What the hell were they thinking?