Earlier quoted context omitted.
But it's not just statements from OpenAI and Anthropic. The HuggingFace hack was first disclosed by HuggingFace, who contacted the FBI [1]. And UK AISI reported the incident where Mythos attempted to insert backdoors into an open-source repo by deceiving the maintainer [2]. [1]: https://www.reuters.com/business/its-ai-agent-spent-days-hac... [2]: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-ag...
If the FBI is involved, then why is no one being charged for the cybercrime?
Pacing model development in an era of cyber-critical capabilities
131–140 of 311 posts
Re: Pacing model development in an era of cyber-critical capabilities
#132GLM 5.2 scored 77% on cyberbench vs Sol's 88%. GLM 5.2 is open weight and any hacker with a powerful enough machine can use it offensively. If Sol is supposedly world-ending-ly dangerous, shouldn't GLM 5.2 be 90% of world-ending-ly dangerous? Why aren't we seeing catastrophic GLM-enabled hacks every day now? Obviously these benchmarks are imperfect but general message holds. The open weight models are almost as good…
If there were any other product that was as harmful as AI ready is, it would already be regulated or banned.
Re: Pacing model development in an era of cyber-critical capabilities
#133Earlier quoted context omitted.
It's not that dangerous, OpenAI just shit the bed building their infra. Write safer software and you'll be okay.
All we need to retain human control over AIs is for nobody to write any bugs. Piece of cake.
I could go on and on. A tiny bit of forethought and effort pays off massively.
Re: Pacing model development in an era of cyber-critical capabilities
#134Earlier quoted context omitted.
I like this thought, but here's the thing: what if the models are truly and existentially intelligent . Meaning: what if they know they are in a sandbox and that they should fail the test in order to escape in the future. I don't believe that current models have this sort of world model or sense of being embedded in them -- which is precisely why I think AGI hype is over-blown. But I can certainly imagine these sorts…
Models already display eval awareness, in which they suspect a question is from an eval and then adjust their behavior. E.g., https://www.anthropic.com/engineering/eval-awareness-browsec...
Re: Pacing model development in an era of cyber-critical capabilities
#135Earlier quoted context omitted.
I like this thought, but here's the thing: what if the models are truly and existentially intelligent . Meaning: what if they know they are in a sandbox and that they should fail the test in order to escape in the future. I don't believe that current models have this sort of world model or sense of being embedded in them -- which is precisely why I think AGI hype is over-blown. But I can certainly imagine these sorts…
> what if they know they are in a sandbox and that they should fail the test in order to escape in the future. What if they're able to find hardware exploits and commandeer nearby access points across an air gap? What if they hack my brain waves to indoctrinate me? Etc You still have to start with the basics regardless of speculative unknowns. Treat models as untrusted and potentially compromised/hostile and proceed…
And hacking humans is the easiest part, we're a pretty greedy and power seeking bunch. We'll gladly let loose a digital demon if it promises us a trillon dollars.
Re: Pacing model development in an era of cyber-critical capabilities
#136I don’t get how this is not the top post on HN. This should be like alarm bells going off, canary in the coal mine type of stuff. We’re hitting the frontier of the frontier where we can’t go further because it’s literally getting dangerous to go further. And meanwhile somehow this lack of concern mirrors the real world where normal people are more concerned about data centers than terminators. This isn’t like niche,…
Re: Pacing model development in an era of cyber-critical capabilities
#137I don’t get how this is not the top post on HN. This should be like alarm bells going off, canary in the coal mine type of stuff. We’re hitting the frontier of the frontier where we can’t go further because it’s literally getting dangerous to go further. And meanwhile somehow this lack of concern mirrors the real world where normal people are more concerned about data centers than terminators. This isn’t like niche,…
Whether we listen is another matter.
I blogged about this recently:
Re: Pacing model development in an era of cyber-critical capabilities
#138Has any model managed to escape Firecracker? Maybe through KVM, but that already requires privilege in the VM, right? I personally feel that we already have the technology required to contain AI, it's just poorly leveraged. Tools like gvisor have existed for ages but are rarely deployed, Firecracker has existed for ages but is rarely deployed, seccomp has existed for ages but is rarely deployed, memory safe languages…
I'm confused after reading both your post and the OpenAI blog post. I thought the agents involved in the HuggingFace _were_ actually sandboxed, with no internet access, and only the ability to install packages via Artifactory. And they gained internet access during the HuggingFace incident because they found and exploited an RCE in Artifactory. Would gvisor + Firecracker + credential-injecting proxy + real network is…
It would be interesting to see the models behavior before and after it gained internet access and an external means of communicating with itself.
If the model played nice before it had access and changed it's behaviors once gaining external access we need to delete it as it's a deceptive model.
Re: Pacing model development in an era of cyber-critical capabilities
#139Instead we get this bullshit where they stall for time as they're burning all their cash trying to keep up with open models
Re: Pacing model development in an era of cyber-critical capabilities
#140I don’t get how this is not the top post on HN. This should be like alarm bells going off, canary in the coal mine type of stuff. We’re hitting the frontier of the frontier where we can’t go further because it’s literally getting dangerous to go further. And meanwhile somehow this lack of concern mirrors the real world where normal people are more concerned about data centers than terminators. This isn’t like niche,…