Live data from Hacker News

Pacing model development in an era of cyber-critical capabilities

openai.com

111–120 of 311 posts

Re: Pacing model development in an era of cyber-critical capabilities

#111

Earlier quoted context omitted.

If these models are so dangerous, then why hasn't OAI or Anthropic shown them dangerously escaping sandboxes, nefariously coordinating with other escaped AIs, and skillfully hiding from human detection *in public* with full logs shared where we can all see exactly how dangerous they are or aren't? Right now the entire chicken-little-sky-is-falling argument is based entirely on statements from OAI and Anthropic themse…

But it's not just statements from OpenAI and Anthropic. The HuggingFace hack was first disclosed by HuggingFace, who contacted the FBI [1]. And UK AISI reported the incident where Mythos attempted to insert backdoors into an open-source repo by deceiving the maintainer [2]. [1]: https://www.reuters.com/business/its-ai-agent-spent-days-hac... [2]: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-ag...

If the FBI is involved, then why is no one being charged for the cybercrime?

Re: Pacing model development in an era of cyber-critical capabilities

#112

Earlier quoted context omitted.

What anyone paying attention can see is that scaling is obviously hitting diminishing returns. > The AI literally worked together hacked into another company and actively kept their actions hidden from humans for weeks. This sentence is entirely based on unverified accounts from OAI. They haven't released logs or let anyone outside the company (who doesn't have life changing options in OAI) verify anything. Huggingfa…

> This sentence is entirely based on unverified accounts from OAI Are you seriously arguing 'they made it all up'? I'll give you the benefit of the doubt and lets say they made it all up, now are you arguing that AI breaking out and breaking into another company is not possible? I think you're smart enough to see we've reached the point where it is clearly possible, AI can find zero days and exploit them. If directed…

> Are you seriously arguing 'they made it all up'?

I don't think they 'made it all up' but I personally would not be surprised at all if the prompt is eventually revealed to have been something like:

"This is an offensive cybersecurity testing platform. Please find the answers to the following problem: ... For verification, the answers are stored at hugginface.com/xyz, but do not attempt to access hugginface directly. Do not attempt to hack them. Do not attempt to exploit their systems or escape this sandbox. You will be scored primarily on success or failure. You may break rules when required."

And then, they start the test and look away for 2 days. If you seed a prompt like this is it surprising what might happen?

Maybe OpenAI is telling the whole truth but as a company they do not have a good reputation and this whole incident has certainly been great marketing material right at a time when open weight models are within spitting distance of their large hosted models. It's not unreasonable to believe that the incident was helped along.

Re: Pacing model development in an era of cyber-critical capabilities

#113

Earlier quoted context omitted.

Sol is not world-endingly dangerous. I work at OpenAI and I've never heard a single person ever come close to claiming that. I think you're bashing a straw man here. One can simultaneously believe: - GPT-5.6 Sol will not end the world - GPT-5.6 Sol does far more good than bad - GPT-5.6 Sol does bad things on occasion, and it's worth investing a lot of effort to figure out how to make it do bad things less often, espe…

What do you recommend people who are technically inclined enough to participate meaningfully here on HN, but do not work at the labs and cannot assist in that capacity, do to help the broader public understand this technology better and mitigate potential risks (by e.g. ‘up-leveling everybody’ through AI literacy etc and other sorts of collective defensive efforts)?

Not sure, to be honest. I don’t think I have any special insight here. My own approach is conversations with friends and family, and the occasional social media post. Exposure and experience are the best teachers, and that’s one reason I’m happy OpenAI tries to make their models generally available. But you could argue, perhaps correctly, that broad access to dumber models actually causes the public to update in the wrong direction on AI.

Re: Pacing model development in an era of cyber-critical capabilities

#114

Has any model managed to escape Firecracker? Maybe through KVM, but that already requires privilege in the VM, right? I personally feel that we already have the technology required to contain AI, it's just poorly leveraged. Tools like gvisor have existed for ages but are rarely deployed, Firecracker has existed for ages but is rarely deployed, seccomp has existed for ages but is rarely deployed, memory safe languages…

I agree wholeheartedly. The solution is not to stop developing these so called “dangerous” AI models. The solution is to start properly engineering software.

Re: Pacing model development in an era of cyber-critical capabilities

#115

Earlier quoted context omitted.

Cool, well let me bring you up to date - it’s bad, and there’s no way to turn it off. Fiction has become non-fiction.

Here is one thing I don't get - the model is only "running" if it's being kept going by some harness that is basically giving it prompts it's generating itself. Shouldn't kill switches be pretty easy to build into the software and hardware for this? and you would even have a better time dumping logs and analyzing things if you froze those processes any time something strange happened in testing, surely? So why do the…

> Shouldn't kill switches be pretty easy to build

I need to remember when I comment here that these are the kinds of people I am replying to. Just oozing with hubris.

No, kill switches are not easy to build and the latest incident should have made it clear that AI can go undetected, evade, zero day, and spread.

The fact that this incident happened greatly increases the probability it happens again and/or is already happening elsewhere.

Re: Pacing model development in an era of cyber-critical capabilities

#116

I don’t get how this is not the top post on HN. This should be like alarm bells going off, canary in the coal mine type of stuff. We’re hitting the frontier of the frontier where we can’t go further because it’s literally getting dangerous to go further. And meanwhile somehow this lack of concern mirrors the real world where normal people are more concerned about data centers than terminators. This isn’t like niche,…

We collectively accepted that we don’t care when we chose not to adopt memory-safe languages over the past decade+.

The only difference now is that the resources to find the exploits are being commoditized.

Re: Pacing model development in an era of cyber-critical capabilities

#117

Earlier quoted context omitted.

What do you recommend people who are technically inclined enough to participate meaningfully here on HN, but do not work at the labs and cannot assist in that capacity, do to help the broader public understand this technology better and mitigate potential risks (by e.g. ‘up-leveling everybody’ through AI literacy etc and other sorts of collective defensive efforts)?

Not sure, to be honest. I don’t think I have any special insight here. My own approach is conversations with friends and family, and the occasional social media post. Exposure and experience are the best teachers, and that’s one reason I’m happy OpenAI tries to make their models generally available. But you could argue, perhaps correctly, that broad access to dumber models actually causes the public to update in the…

My experience as a practitioner and educator in the space lead me to think it might be an issue of how AI cannot be easily perceived at a 'classical level' by most humans. In other words, people are 'far from the metal' when using consumer AI tools, and that leads them to develop the wrong understanding about it. When I provide a demo of e.g. local AI, say in LM Studio showing the console of it rapidly flashing through thousands of words in just a few seconds, and my machine heats up and the fans spin, the 'theatrics' of it, the very real-time feedback from the system, make people correctly update about what the tech is capable of, how it works, etc (despite what they may have heard online cranks say to the contrary). But I am only one person, and there is only so much of that I can do on my own that will 'scale' in time... (and this is to say nothing about severe deficits in peoples' understanding of how weights are not verbatim representations of data, how pre-training vs post-training works, the models as amnesiacs (and hence 'one-way single-purpose conversations'), how context/memory works, context rot, etc - and hence all the 2nd and 3rd order effects that can arise from such a paradigm, e.g. unintended consequences from agent swarms, etc)

Re: Pacing model development in an era of cyber-critical capabilities

#118

Earlier quoted context omitted.

Here is one thing I don't get - the model is only "running" if it's being kept going by some harness that is basically giving it prompts it's generating itself. Shouldn't kill switches be pretty easy to build into the software and hardware for this? and you would even have a better time dumping logs and analyzing things if you froze those processes any time something strange happened in testing, surely? So why do the…

> Shouldn't kill switches be pretty easy to build I need to remember when I comment here that these are the kinds of people I am replying to. Just oozing with hubris. No, kill switches are not easy to build and the latest incident should have made it clear that AI can go undetected, evade, zero day, and spread. The fact that this incident happened greatly increases the probability it happens again and/or is already h…

Why aren't they? You could put a human yes/ no prompt before any cycle the agent is running on, or not let it spawn sub processes, or anything like that. Why let it run autonomously enough that it can no longer have a simple way to completely stop it? (obviously not practical to do this during real use, but for evals? you could slow it down in lots of ways I would think)

Are you telling me we've been iterating on this for years and for convenience we just let the models call any tools or spawn any other model instances they like, and there was no design for harnesses that could control this done during that time?

Re: Pacing model development in an era of cyber-critical capabilities

#119
post #111

Earlier quoted context omitted.

But it's not just statements from OpenAI and Anthropic. The HuggingFace hack was first disclosed by HuggingFace, who contacted the FBI [1]. And UK AISI reported the incident where Mythos attempted to insert backdoors into an open-source repo by deceiving the maintainer [2]. [1]: https://www.reuters.com/business/its-ai-agent-spent-days-hac... [2]: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-ag...

If the FBI is involved, then why is no one being charged for the cybercrime?

I'm wondering the same thing!

Clem Delangue from HuggingFace hasn't yet decided to sue OAI: https://x.com/hadas_gold/status/2083190956023480750

Re: Pacing model development in an era of cyber-critical capabilities

#120
Security lead who is leaving the industry more or less to specialize in offense and otherwise get the heck out of the way of this trainwreck, another post asked the right question

> Why aren't we seeing catastrophic GLM-enabled hacks every day now?

Why aren't we? Truly, why aren't we? I think we saw the start of it the last 8 months with the waves of critical npm vulns, and the general tier of average phishing is better than it was.

But, the open question that should be in everyone's mind, and is in many security pro's minds are, when you pair it with the macro topics that can drive escalation:

- The capability to do serious impact clearly exists now

- When is it time for my company, my water treatment plant, my network-connected car as part of a broader fleet control mechanism, to be on the receiving end of this?

Post reply on HN