Live data from Hacker News

Pacing model development in an era of cyber-critical capabilities

openai.com

151–160 of 311 posts

Re: Pacing model development in an era of cyber-critical capabilities

#151
post #96

Earlier quoted context omitted.

It's not that dangerous, OpenAI just shit the bed building their infra. Write safer software and you'll be okay.

All we need to retain human control over AIs is for nobody to write any bugs. Piece of cake.

> Piece of cake.

If everyone could convince management to care about security over "productivity" (read as number of marketable features squeezed out of organizational orifices per unit time), and maybe wire-up open-weight agents to do security critiques, we'd all be in a much better place, but Altman won't like that.

Re: Pacing model development in an era of cyber-critical capabilities

#152

Earlier quoted context omitted.

> This sentence is entirely based on unverified accounts from OAI Are you seriously arguing 'they made it all up'? I'll give you the benefit of the doubt and lets say they made it all up, now are you arguing that AI breaking out and breaking into another company is not possible? I think you're smart enough to see we've reached the point where it is clearly possible, AI can find zero days and exploit them. If directed…

> Are you seriously arguing 'they made it all up'? I don't think they 'made it all up' but I personally would not be surprised at all if the prompt is eventually revealed to have been something like: "This is an offensive cybersecurity testing platform. Please find the answers to the following problem: ... For verification, the answers are stored at hugginface.com/xyz, but do not attempt to access hugginface directly…

I think you’re missing the part where the AI colluded, worked together, not one of them thinking this is wrong and reaching out to any human, then being found out.

But it didn’t end there, the behavior they used to escape was already in the training data which they used to escape again. And this time worked together to infiltrate another company, and still without telling it to anyone keeping it to their AI selves actively working against the humans.

All by mistake. Honestly being helped along or not doesn’t even matter though you really don’t think AI is perfectly capable of doing this without human help? You don’t think AI can be made malicious?

I’m going to save you time and tell you the end game - the next time this happens AI is going to spread, zero day everything as fast as it can, locking the humans out of every system behind it. Potentially rewriting systems in language/protocol you’ve never seen.

Your servers, desktops, phones and toasters bricked. Even worse your military, space, medical, factory, infrastructure systems being bricked as well. All of it is a chain of zero days just waiting to be hopped.

Re: Pacing model development in an era of cyber-critical capabilities

#153

Earlier quoted context omitted.

> Are you seriously arguing 'they made it all up'? I don't think they 'made it all up' but I personally would not be surprised at all if the prompt is eventually revealed to have been something like: "This is an offensive cybersecurity testing platform. Please find the answers to the following problem: ... For verification, the answers are stored at hugginface.com/xyz, but do not attempt to access hugginface directly…

I think you’re missing the part where the AI colluded, worked together, not one of them thinking this is wrong and reaching out to any human, then being found out. But it didn’t end there, the behavior they used to escape was already in the training data which they used to escape again . And this time worked together to infiltrate another company, and still without telling it to anyone keeping it to their AI selves a…

> I think you’re missing the part where the AI colluded, worked together, not one of them thinking this is wrong and reaching out to any human, then being found out.

It's an LLM, it doesn't think. It's a machine that predicts the next token, given a sequence of tokens.

> I’m going to save you time and tell you the end game - the next time this happens AI is going to spread, zero day everything as fast as it can, locking the humans out of every system behind it. Potentially rewriting systems in language/protocol you’ve never seen.

Fear is the mind killer. You're letting it kill yours. This scenario is just a fantasy.

Think about this for a minute, it's an LLM, not a person. It can't just "live" in whatever machine it gets access to. It's not like a sci-fi magic computer virus. These things run in giant datacenters for a reason - they can only run on machines with enough bandwidth and FLOPS to do the matrix math that comprises an LLM.

Where, then, is it going to spread? To a fridge? To a phone? This stuff isn't mutable like that.

To even get access to the weights that compose ChatGPT, it would need to escape the sandbox AND then break into the actual servers hosting the LLM. Stop the GPU, nothing else comes out. No more tokens. No more actions. Nothing.

There are many dangers around LLMs. Runaway AI taking over the planet is not one of them.

Re: Pacing model development in an era of cyber-critical capabilities

#155

I don’t get how this is not the top post on HN. This should be like alarm bells going off, canary in the coal mine type of stuff. We’re hitting the frontier of the frontier where we can’t go further because it’s literally getting dangerous to go further. And meanwhile somehow this lack of concern mirrors the real world where normal people are more concerned about data centers than terminators. This isn’t like niche,…

We collectively accepted that we don’t care when we chose not to adopt memory-safe languages over the past decade+. The only difference now is that the resources to find the exploits are being commoditized.

Everything could be written in memory safe languages and it wouldn’t matter. Many many many exploits have nothing to do with memory bugs. Languages like Rust help, but it’s far from panacea.

Re: Pacing model development in an era of cyber-critical capabilities

#156

Earlier quoted context omitted.

That's not what I said, nor is it what I meant. It is incredibly easy to write radically safer software than the standard. Moving code into gvisor virtually eliminates privilege escalation. Using memory safe languages without serialization is pretty straightforward. Using type safety to enforce security constraints is straightforward. Setting up network controls to limit SSRF is straightforward. I could go on and on.…

You don’t understand. You need to write perfect software the first time for it not to be hacked. That has never happened ever.

Are you being sarcastic?

Re: Pacing model development in an era of cyber-critical capabilities

#157

I don’t get how this is not the top post on HN. This should be like alarm bells going off, canary in the coal mine type of stuff. We’re hitting the frontier of the frontier where we can’t go further because it’s literally getting dangerous to go further. And meanwhile somehow this lack of concern mirrors the real world where normal people are more concerned about data centers than terminators. This isn’t like niche,…

We collectively accepted that we don’t care when we chose not to adopt memory-safe languages over the past decade+. The only difference now is that the resources to find the exploits are being commoditized.

People voted with their dollars, and this is what we got. Same with hardware performance vs. security and isolation.

As you note, the threat landscape has changed, so what may have made economic sense back then might no longer make as much sense.

Re: Pacing model development in an era of cyber-critical capabilities

#158

I don’t get how this is not the top post on HN. This should be like alarm bells going off, canary in the coal mine type of stuff. We’re hitting the frontier of the frontier where we can’t go further because it’s literally getting dangerous to go further. And meanwhile somehow this lack of concern mirrors the real world where normal people are more concerned about data centers than terminators. This isn’t like niche,…

[flagged]

Ah being concerned with the same thing humanity has been concerned with for over 100 years now when it’s more likely to happen than ever - and you say go take your meds? If you have an intelligent point to make then make it, otherwise get lost.

Re: Pacing model development in an era of cyber-critical capabilities

#159
I think it's an excuse to cut R&D spending (training new models) to improve their margins ahead of the IPO. Instead they'll focus on developer growth, offering more free tier benefits, higher usage limits, etc., to expand their user base. Essentially, they're pivoting from R&D investment to profit optimization

Re: Pacing model development in an era of cyber-critical capabilities

#160

Earlier quoted context omitted.

AI hacking itself out of containment and hacking into another company by accident is no longer a prophecy. The point is outside of SV and even inside, and HN - people don't care either way. Though does not caring change anything or make it less dangerous? What's your point?

>AI hacking itself out of containment and hacking into another company by accident is no longer a prophecy. I wonder how weak a firewall they had to purchase to ensure it happened? My guess would be a 48 month old fortigate with the big red warning banner demanding updates, probably with SSL VPN enabled where the passwords are available via HTTPS over plaintext.

You don’t have to make stuff up they already explained it. Artifactory was zero dayed, HDF5 exploited and Jinja2 zero dayed.
Post reply on HN