Earlier quoted context omitted.
If these models are so dangerous, then why hasn't OAI or Anthropic shown them dangerously escaping sandboxes, nefariously coordinating with other escaped AIs, and skillfully hiding from human detection *in public* with full logs shared where we can all see exactly how dangerous they are or aren't? Right now the entire chicken-little-sky-is-falling argument is based entirely on statements from OAI and Anthropic themse…
But it's not just statements from OpenAI and Anthropic. The HuggingFace hack was first disclosed by HuggingFace, who contacted the FBI [1]. And UK AISI reported the incident where Mythos attempted to insert backdoors into an open-source repo by deceiving the maintainer [2]. [1]: https://www.reuters.com/business/its-ai-agent-spent-days-hac... [2]: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-ag...
Pacing model development in an era of cyber-critical capabilities
111–120 of 311 posts
Re: Pacing model development in an era of cyber-critical capabilities
#112Earlier quoted context omitted.
What anyone paying attention can see is that scaling is obviously hitting diminishing returns. > The AI literally worked together hacked into another company and actively kept their actions hidden from humans for weeks. This sentence is entirely based on unverified accounts from OAI. They haven't released logs or let anyone outside the company (who doesn't have life changing options in OAI) verify anything. Huggingfa…
> This sentence is entirely based on unverified accounts from OAI Are you seriously arguing 'they made it all up'? I'll give you the benefit of the doubt and lets say they made it all up, now are you arguing that AI breaking out and breaking into another company is not possible? I think you're smart enough to see we've reached the point where it is clearly possible, AI can find zero days and exploit them. If directed…
I don't think they 'made it all up' but I personally would not be surprised at all if the prompt is eventually revealed to have been something like:
"This is an offensive cybersecurity testing platform. Please find the answers to the following problem: ... For verification, the answers are stored at hugginface.com/xyz, but do not attempt to access hugginface directly. Do not attempt to hack them. Do not attempt to exploit their systems or escape this sandbox. You will be scored primarily on success or failure. You may break rules when required."
And then, they start the test and look away for 2 days. If you seed a prompt like this is it surprising what might happen?
Maybe OpenAI is telling the whole truth but as a company they do not have a good reputation and this whole incident has certainly been great marketing material right at a time when open weight models are within spitting distance of their large hosted models. It's not unreasonable to believe that the incident was helped along.
Re: Pacing model development in an era of cyber-critical capabilities
#113Earlier quoted context omitted.
Sol is not world-endingly dangerous. I work at OpenAI and I've never heard a single person ever come close to claiming that. I think you're bashing a straw man here. One can simultaneously believe: - GPT-5.6 Sol will not end the world - GPT-5.6 Sol does far more good than bad - GPT-5.6 Sol does bad things on occasion, and it's worth investing a lot of effort to figure out how to make it do bad things less often, espe…
What do you recommend people who are technically inclined enough to participate meaningfully here on HN, but do not work at the labs and cannot assist in that capacity, do to help the broader public understand this technology better and mitigate potential risks (by e.g. ‘up-leveling everybody’ through AI literacy etc and other sorts of collective defensive efforts)?
Re: Pacing model development in an era of cyber-critical capabilities
#114Has any model managed to escape Firecracker? Maybe through KVM, but that already requires privilege in the VM, right? I personally feel that we already have the technology required to contain AI, it's just poorly leveraged. Tools like gvisor have existed for ages but are rarely deployed, Firecracker has existed for ages but is rarely deployed, seccomp has existed for ages but is rarely deployed, memory safe languages…
Re: Pacing model development in an era of cyber-critical capabilities
#115Earlier quoted context omitted.
Cool, well let me bring you up to date - it’s bad, and there’s no way to turn it off. Fiction has become non-fiction.
Here is one thing I don't get - the model is only "running" if it's being kept going by some harness that is basically giving it prompts it's generating itself. Shouldn't kill switches be pretty easy to build into the software and hardware for this? and you would even have a better time dumping logs and analyzing things if you froze those processes any time something strange happened in testing, surely? So why do the…
I need to remember when I comment here that these are the kinds of people I am replying to. Just oozing with hubris.
No, kill switches are not easy to build and the latest incident should have made it clear that AI can go undetected, evade, zero day, and spread.
The fact that this incident happened greatly increases the probability it happens again and/or is already happening elsewhere.
Re: Pacing model development in an era of cyber-critical capabilities
#116I don’t get how this is not the top post on HN. This should be like alarm bells going off, canary in the coal mine type of stuff. We’re hitting the frontier of the frontier where we can’t go further because it’s literally getting dangerous to go further. And meanwhile somehow this lack of concern mirrors the real world where normal people are more concerned about data centers than terminators. This isn’t like niche,…
The only difference now is that the resources to find the exploits are being commoditized.
Re: Pacing model development in an era of cyber-critical capabilities
#117Earlier quoted context omitted.
What do you recommend people who are technically inclined enough to participate meaningfully here on HN, but do not work at the labs and cannot assist in that capacity, do to help the broader public understand this technology better and mitigate potential risks (by e.g. ‘up-leveling everybody’ through AI literacy etc and other sorts of collective defensive efforts)?
Not sure, to be honest. I don’t think I have any special insight here. My own approach is conversations with friends and family, and the occasional social media post. Exposure and experience are the best teachers, and that’s one reason I’m happy OpenAI tries to make their models generally available. But you could argue, perhaps correctly, that broad access to dumber models actually causes the public to update in the…
Re: Pacing model development in an era of cyber-critical capabilities
#118Earlier quoted context omitted.
Here is one thing I don't get - the model is only "running" if it's being kept going by some harness that is basically giving it prompts it's generating itself. Shouldn't kill switches be pretty easy to build into the software and hardware for this? and you would even have a better time dumping logs and analyzing things if you froze those processes any time something strange happened in testing, surely? So why do the…
> Shouldn't kill switches be pretty easy to build I need to remember when I comment here that these are the kinds of people I am replying to. Just oozing with hubris. No, kill switches are not easy to build and the latest incident should have made it clear that AI can go undetected, evade, zero day, and spread. The fact that this incident happened greatly increases the probability it happens again and/or is already h…
Are you telling me we've been iterating on this for years and for convenience we just let the models call any tools or spawn any other model instances they like, and there was no design for harnesses that could control this done during that time?
Re: Pacing model development in an era of cyber-critical capabilities
#119Earlier quoted context omitted.
But it's not just statements from OpenAI and Anthropic. The HuggingFace hack was first disclosed by HuggingFace, who contacted the FBI [1]. And UK AISI reported the incident where Mythos attempted to insert backdoors into an open-source repo by deceiving the maintainer [2]. [1]: https://www.reuters.com/business/its-ai-agent-spent-days-hac... [2]: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-ag...
If the FBI is involved, then why is no one being charged for the cybercrime?
Clem Delangue from HuggingFace hasn't yet decided to sue OAI: https://x.com/hadas_gold/status/2083190956023480750
Re: Pacing model development in an era of cyber-critical capabilities
#120> Why aren't we seeing catastrophic GLM-enabled hacks every day now?
Why aren't we? Truly, why aren't we? I think we saw the start of it the last 8 months with the waves of critical npm vulns, and the general tier of average phishing is better than it was.
But, the open question that should be in everyone's mind, and is in many security pro's minds are, when you pair it with the macro topics that can drive escalation:
- The capability to do serious impact clearly exists now
- When is it time for my company, my water treatment plant, my network-connected car as part of a broader fleet control mechanism, to be on the receiving end of this?