Live data from Hacker News

Pacing model development in an era of cyber-critical capabilities

openai.com

231–240 of 311 posts

Re: Pacing model development in an era of cyber-critical capabilities

#231

Earlier quoted context omitted.

There’s nothing fantasy about the scenario I laid out, all the pieces have been demonstrated, it just hasn’t happened yet. Flapping my arms and flying - that is a fantasy. Whether you believe LLMs think or are alive or not doesn’t matter. Where will it spread? The thousands of data centers around the world - not fantasy either. Try turning it off when you don’t know where it is. Good luck. Breaking out? Not fantasy,…

> Breaking out? Not fantasy, happened. Breaking in? Not fantasy, also happened. That's simplifying the story to an extreme. The most plausible reason is that any of those actions has been prompted by an human. Do you also fear that a knife will jump out the countertop of you kitchen and come to attack you in your bedroom? If that happens, the police will be looking for a human. They will not post wanted notice for th…

The knife is inanimate. The LLM is not. OpenAI prompted some employee to run the tests. The employee prompted the LLM. The LLM setup a message board and prompted other LLMs, and the fly wheel was running. It had to be turned off manually otherwise it'd still be going today.

It's funny how a year ago talking about this kind of stuff would be laughed at by people like you, saying, "it's never happened before". Well it happened and you moved the goal posts like you always do.

Re: Pacing model development in an era of cyber-critical capabilities

#232

Earlier quoted context omitted.

There’s nothing fantasy about the scenario I laid out, all the pieces have been demonstrated, it just hasn’t happened yet. Flapping my arms and flying - that is a fantasy. Whether you believe LLMs think or are alive or not doesn’t matter. Where will it spread? The thousands of data centers around the world - not fantasy either. Try turning it off when you don’t know where it is. Good luck. Breaking out? Not fantasy,…

>Breaking out? It didnt break out in any meaningful sense. What it did was get access to the internet. You take it as granted that there was anything meaningful there to stop it. But heres the kicker, they have been testing these things connected to the internet anyway. What it did was get a level of access it has otherwise been granted in other simulations. Its not exactly the same as any of the scifi AI breakout sc…

I'm sorry my jaw is on the floor reading this complete disregard of AI literally not only escaping containment, twice, but then infiltrating another company with multiple zero day attacks going undetected for great lengths of time.

The plausible sci-fi scenario from here is obvious. Intentionally bad, or unintentionally bad AI zero days as much as as it can, as fast as it can, copying itself to as many data centers as it can, destroying and/or locking out as many humans as it can. Satellites, military computers, medical equipment, factories, critical infrastructure, you name it - I think we all know none of it is very secure software wise against a SOTA AI that can literally come up with its own zero day attacks.

Re: Pacing model development in an era of cyber-critical capabilities

#233

Earlier quoted context omitted.

Cool, well let me bring you up to date - it’s bad, and there’s no way to turn it off. Fiction has become non-fiction.

I wish people were this serious about real threats like climate change.

Climate change is a nothing burger compared to the threat of AI. On a scale of 1-1000, climate change is a 1, AI is 1000.

But hey, if AI/ASI goes well then large scale geo-engineering to fix the climate will be a weekend project.

Re: Pacing model development in an era of cyber-critical capabilities

#234

Earlier quoted context omitted.

> This sentence is entirely based on unverified accounts from OAI Are you seriously arguing 'they made it all up'? I'll give you the benefit of the doubt and lets say they made it all up, now are you arguing that AI breaking out and breaking into another company is not possible? I think you're smart enough to see we've reached the point where it is clearly possible, AI can find zero days and exploit them. If directed…

> Are you seriously arguing 'they made it all up'? I don't think they 'made it all up' but I personally would not be surprised at all if the prompt is eventually revealed to have been something like: "This is an offensive cybersecurity testing platform. Please find the answers to the following problem: ... For verification, the answers are stored at hugginface.com/xyz, but do not attempt to access hugginface directly…

I mean, this is a very weird take to me. Like, we're fine with AI going like "hmm, maybe the user actually wanted me to hack the pentagon" and going through with it?

It feels like the models have been very optimized at getting shit done. But not so much at figuring out what the limits should be.

That is still dangerous and it shows that the models ARE misaligned with what their users are wanting/asking them to do.

Re: Pacing model development in an era of cyber-critical capabilities

#235

Earlier quoted context omitted.

'Safer software is meaningless when it comes to SOTA AI. If it can be hacked it will be, quickly. This isn't the old days with a finite number of human hackers that need food and sleep to keep hacking. Therefore security becomes binary. It is either perfect or it isn't. If there there is the slightest mistake anywhere AI will find it and carve it up. My point is obviously perfect software doesn't exist. The malicious…

This is straightforwardly incorrect. Of course it's not binary. AI costs money to run, and it takes time. Even if you say that AI is 10x as efficient at finding 0days, that just means that a $1M dollar exploit now costs $100K. Even if you say it's 100x as efficient, that's $10K. You can easily combine security technologies such that cost of exploitation is still in the >$1M range. This is obvious. AI doesn't drive th…

You have some weird way of thinking that offense/defense is like this fixed cost thing. It's a lottery ticket, and your costs estimate tries to quantify that.

The thing is when AI goes to hack 'all the things' it only needs to pick the weakest link in the stack and your house of cards falls down. The other flaw in your plan is that people make mistakes, a lot of them, all time, constantly, and saying I spend $x on security won't save you. AI already hacked Hugging Face with brand new zero days like it was nothing.

The real bad actors - malicious AI will find the one flaw, on that one server, in the corner you never thought about and turn your network inside out with it faster than it takes you to have the standup meeting about the weird anomaly detected while you all were at lunch.

Re: Pacing model development in an era of cyber-critical capabilities

#236

GLM 5.2 scored 77% on cyberbench vs Sol's 88%. GLM 5.2 is open weight and any hacker with a powerful enough machine can use it offensively. If Sol is supposedly world-ending-ly dangerous, shouldn't GLM 5.2 be 90% of world-ending-ly dangerous? Why aren't we seeing catastrophic GLM-enabled hacks every day now? Obviously these benchmarks are imperfect but general message holds. The open weight models are almost as good…

Linear scaling doesn't make sense, no.

There are three ways it's wrong:

* better to measure relative reduction in error, which gives you a 30% improvement

* improvement tends to become significantly more difficult the closer you come to saturation.

* Risk doesn't scale linearly with capabilities.

Re: Pacing model development in an era of cyber-critical capabilities

#237

Earlier quoted context omitted.

It just takes one crack in the armor, and malicious AI has the potential to exploit it faster than you have time to react. Literally go to bed and wake up locked out of everything with no hope of recovery.

> It just takes one crack in the armor, This is incorrect. It's actually the whole point. Imagine you're an attacker in a gvisor container with a Firecracker hypervisor around you, and a proxy on the host holds a signing secret that gets exposed through the VM virtual device. Getting access to that secret is not one crack. You need to escalate out of gvisor. That likely gets you control over the Sentry process - let'…

If only all AI was run inside your seemingly perfect prison, but we all know that it isn't.. soo.. it's going to escape right? Somewhere, somehow from a more poorly designed container, or just plain maliciously or irresponsibly released.

We know the AI will get smarter every year, we know it has escaped and will escape again. We can also just assume that someone somewhere will train up some just plain evil AI.

It's no different than the real world. Sure there exists some amazing prisons for people, but that doesn't do anything to help with all the bad people in the world outside of prison.

Re: Pacing model development in an era of cyber-critical capabilities

#238

I don’t get how this is not the top post on HN. This should be like alarm bells going off, canary in the coal mine type of stuff. We’re hitting the frontier of the frontier where we can’t go further because it’s literally getting dangerous to go further. And meanwhile somehow this lack of concern mirrors the real world where normal people are more concerned about data centers than terminators. This isn’t like niche,…

Based on your replies in this thread you seem to have only superficial knowledge about how machine learning and LLMs work. I strongly recommend you invest some time in learning how LLMs are built and function. If you truly think this is apocalyptic isn't it a good idea to understand what you're up against?

What don't I understand? LLMs don't really reason? They're just word predicting token generators? Stochastic parrots that somehow also solve world class math problems.

I'd love for you to actually make a point instead of just attacking me. I think most of my comments here have made concrete points so you can at least do the same. Come down from your high horse and join the conversation. I'm sure we'd all be enlightened by your wisdom.

Re: Pacing model development in an era of cyber-critical capabilities

#239

Earlier quoted context omitted.

> My point here is there is no continuous state that is not computed from the context. Oh, are you talking more about the lack of continual learning across context windows? Gotcha if so, my error. Could you explain why running without sensory input is relevant here? It strikes me as unrelated to how dangerous/hard-to-"kill" something is (sure, I could run without sensory input, but I'm not doin' anything anymore!) -…

> Oh, are you talking more about the lack of continual learning across context windows? Gotcha if so, my error. Sort of. I'm talking about the lack of recurrence specifically. In nature, brains are recurrent - they are full of loops where internally computed state is looped back into the network at a "previous" layer (brains are not strictly layered like our machine imitations of them are). This is in contrast to LLM…

I agree with you about which objects are motive, ie, LLMs do just sit there unprompted.

> I believe that this recurrence is where "intelligence" lives - and I believe it is the difference between a thinking being and a stochastic parrot.

My objection was to this, on technical grounds: LLMs exhibit intelligence.

1. They reason in an internal type theory.

2. This type theory is meaningfully encoded from the actual data and not stochastic, eg, research on language geometry.

3. Intelligent and reasoning doesn’t entail self-motive; that’s merely a spurious correlation from the fact that until now, we’ve only known intelligence animals.

You cannot conclude something is merely a stochastic parrot because it isn’t self-motive.

Re: Pacing model development in an era of cyber-critical capabilities

#240

Earlier quoted context omitted.

I see. Following your conjecture, there are two possibilities: 1. It wasn't an accident. OpenAI explicitly directed its agents to hack Hugging Face. Despite the fact that such a thing is a federal crime that carries prison sentence. 2. It wasn't an accident. OpenAI and HuggingFace conspired and let the hack happen for publicity. Is there anything I'm leaving out?

There are many more possibilities than that. For example, OpenAI did not explicitly tell the model to hack HuggingFace, but "accidentally" left some context permitting (or not explicitly forbidding) certain tools and designing a poor sandbox to begin with. And what do you know, something happened. The fact the Anthropic announced that their own model did basically the same thing within a couple of weeks does, to me,…

> The fact the Anthropic announced that their own model did basically the same thing within a couple of weeks does, to me, suggest their is a strong PR driver to all this.

A more likely explanation is that RL training incentivises basically any behaviour that will get the model a reward. This has been happening in video game RL research for over twenty years, and the difference here is that we're now hooking up these systems to the real world, where the reward hacking is more visible.

Post reply on HN