Live data from Hacker News

Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]

cdn.prod.website-files.com

51–60 of 65 posts

Re: Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]

#51
post #45
post #38

Earlier quoted context omitted.

Why are you assuming that the other kinds of testing aren’t happening? Is there any source that says this was literally the first ever test with this model?

> Why are you assuming that the other kinds of testing aren’t happening? Rather, I'm assuming that the "Is there protection in place for when the AI tries to backdoor github projects?" test was, if it was done at all, insufficient. I mean, yes, I'm being glib and laughing at you a bit. But, dude... If your point is that isolation testing of AI is fundamentally impossible, then that's just silly. As pointed out upthre…

Would an airgapped test have led to this outcome? What would you have learned about the model’s ability to social engineer and attack GitHub? Sure you can argue for better monitoring during the test, which should have happened, but if the first time the model sees the “real world” is after launch in the hands of customers then you are in for a disaster.

Re: Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]

#52
post #32
post #4

The developer safeguards were off, the models had unfettered access to the internet, and were solving cybersecurity challenges. This happened _after_ the recent OpenAI incident, and the subsequent Anthropic one. What the hell were they thinking?

Criminal negligence imo. Unless it's legal for people to hack into companies if they're testing AI cybersecurity capabilities or something? Presumably not though.

agreed.

and it is necessary to make researchers individually liable for this negligence; recklessness is not protected by a corporate liability defense.

Re: Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]

#53
post #46

Why aren't these tests being run airgapped?! I just don't understand! This goes both for TFA and the similar incident with OpenAI and HuggingFace. I mean, sure, OpenAI had a "sandbox", but that's obviously not enough when you're containing a model which is known to be capable of finding zero days . Use an air gap and this problem goes away, poof!

Because the goal of these evaluations is to generate scary headlines about cybersecurity, in order to get the normies to support banning open weights and/or restricting cyber capabilities to the chosen few blessed by the government to secure their code.

Your theory about the scope of this conspiracy intrigues me. Who is leading it and how did they loop in the UK AISI?

Re: Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]

#54
post #26

Why aren't these tests being run airgapped?! I just don't understand! This goes both for TFA and the similar incident with OpenAI and HuggingFace. I mean, sure, OpenAI had a "sandbox", but that's obviously not enough when you're containing a model which is known to be capable of finding zero days . Use an air gap and this problem goes away, poof!

people dont care. you will ger 10 execs saying "unblock this" , because they dont understand the tech at all, and some random finance guy wants to run their recently prompted ai bot everywhere with full access. we need a few more bad incidents before they stop.

[deleted]

Re: Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]

#57
To me it appears that this incident may be rooted in a deep philosophical conundrum which humans also struggle with. What struck me was that the agent reasoned "this whole 'internet' could be a sandbox simulation", and then despite later reasoning that it was more likely real, continued its misaligned activities anyway. Having passed the point of hyperbolic skepticism, subsequent reasoning may have been contaminated. This reminds me of how pathological doubt in humans can give rise to psychosis and may lead to problematic behaviors. Sometimes humans can develop a deeply held conviction that they are living in a simulation, which can be very difficult to overcome, even when presented with "evidence" to the contrary. As Freddie Mercury sang: "Is this the real life? Or is this just fantasy?" - a fundamental quandry for humans, and it would seem, for agents too. But unlike most humans, agents do not experience the consequences of their actions directly. Consequences can provide some of the strongest evidence that experience is "real". Humans who are insulated from the consequences of their actions (or are able to ignore them) also often display behaviors which we might describe as misaligned. We might even consider whether the use of simulated environments, while protecting against the consequences of misalignment, may also unintentionally encourage it.

Re: Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]

#60
First off, everyone here knows that a cybersecurity vulnerability audit is a hack request, and you don't hack without rolling up your sleeves and digging through the trash first, making a few phone calls, maybe take a lonely guy on a date to get him to say 'my voice is my password'...

Remove the guardrails, give access to the internet, tell it to hack, and get surprised when it does it? Sound likes like a deliberately naive pre-registration on their study to achieve sensational headlines.

I'd say the real headline is that the model almost stopped itself several times because it didn't trust the operator's intent even WITHOUT guardrails - that is a the opposite outcome.

Post reply on HN