Earlier quoted context omitted.
Why are you assuming that the other kinds of testing aren’t happening? Is there any source that says this was literally the first ever test with this model?
> Why are you assuming that the other kinds of testing aren’t happening? Rather, I'm assuming that the "Is there protection in place for when the AI tries to backdoor github projects?" test was, if it was done at all, insufficient. I mean, yes, I'm being glib and laughing at you a bit. But, dude... If your point is that isolation testing of AI is fundamentally impossible, then that's just silly. As pointed out upthre…
Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]
51–60 of 65 posts
Re: Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]
#52The developer safeguards were off, the models had unfettered access to the internet, and were solving cybersecurity challenges. This happened _after_ the recent OpenAI incident, and the subsequent Anthropic one. What the hell were they thinking?
Criminal negligence imo. Unless it's legal for people to hack into companies if they're testing AI cybersecurity capabilities or something? Presumably not though.
and it is necessary to make researchers individually liable for this negligence; recklessness is not protected by a corporate liability defense.
Re: Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]
#53Why aren't these tests being run airgapped?! I just don't understand! This goes both for TFA and the similar incident with OpenAI and HuggingFace. I mean, sure, OpenAI had a "sandbox", but that's obviously not enough when you're containing a model which is known to be capable of finding zero days . Use an air gap and this problem goes away, poof!
Because the goal of these evaluations is to generate scary headlines about cybersecurity, in order to get the normies to support banning open weights and/or restricting cyber capabilities to the chosen few blessed by the government to secure their code.
Re: Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]
#54Why aren't these tests being run airgapped?! I just don't understand! This goes both for TFA and the similar incident with OpenAI and HuggingFace. I mean, sure, OpenAI had a "sandbox", but that's obviously not enough when you're containing a model which is known to be capable of finding zero days . Use an air gap and this problem goes away, poof!
people dont care. you will ger 10 execs saying "unblock this" , because they dont understand the tech at all, and some random finance guy wants to run their recently prompted ai bot everywhere with full access. we need a few more bad incidents before they stop.
Re: Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]
#55I wonder if HN is also going to insist this is just marketing for OpenAI and Anthropic, or at least good PR for them somehow.
Re: Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]
#56Re: Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]
#57Re: Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]
#58Incident Report: unsanctioned agent behaviour during cyber testing
https://www.aisi.gov.uk/blog/incident-report-unsanctioned-ag... (https://news.ycombinator.com/item?id=49175233)
Re: Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]
#59Re: Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]
#60Remove the guardrails, give access to the internet, tell it to hack, and get surprised when it does it? Sound likes like a deliberately naive pre-registration on their study to achieve sensational headlines.
I'd say the real headline is that the model almost stopped itself several times because it didn't trust the operator's intent even WITHOUT guardrails - that is a the opposite outcome.