Live data from Hacker News

Pacing model development in an era of cyber-critical capabilities

openai.com

121–130 of 311 posts

Re: Pacing model development in an era of cyber-critical capabilities

#121

I don’t get how this is not the top post on HN. This should be like alarm bells going off, canary in the coal mine type of stuff. We’re hitting the frontier of the frontier where we can’t go further because it’s literally getting dangerous to go further. And meanwhile somehow this lack of concern mirrors the real world where normal people are more concerned about data centers than terminators. This isn’t like niche,…

Follow the money: who told you that OpenAI's models autonomously coordinated to hack external systems? What incentives might they have to want you to believe that story? Are there priors which demonstrate them benefiting from telling similar stories, regardless of their factuality?

But to your counterpoint, let's say the story is 100% true, because I agree it is at least plausible. What would the incentive be for the HN audience to believe it? What priors might support their disbelief? For my part, I don't think it's because people lack imagination. I think it's quite rational to question the authenticity and impact of the claims being made. What's worse, believing the story and being wrong, or not believing the story and being wrong?

That said, I agree with you: the impact we're having by not changing course is quite dangerous, the scale is dangerous, the inability to reverse the harms is dangerous, and the lack of collective effort to regulate further damage is dangerous.

It's true, humans are dangerous when trillions are involved. See: climate change.

Re: Pacing model development in an era of cyber-critical capabilities

#122

Earlier quoted context omitted.

Cool, well let me bring you up to date - it’s bad, and there’s no way to turn it off. Fiction has become non-fiction.

Here is one thing I don't get - the model is only "running" if it's being kept going by some harness that is basically giving it prompts it's generating itself. Shouldn't kill switches be pretty easy to build into the software and hardware for this? and you would even have a better time dumping logs and analyzing things if you froze those processes any time something strange happened in testing, surely? So why do the…

You need to read the AI safety stuff from the people that you say are from 2010. There are plenty of good arguments on why kill switches will never work.

AI is already being heavily used in cyber warfare by governments. You think they want easy to break systems when 'enemy' AI will most certainly attack that first? We'll find over time that agentic systems get harder and harder to 'kill' because allowing any AI on the internet that is easy to kill will get it DDOSed.

Also building it into software is nearly useless as AI can write and make software. Just replace and kill your loop with theirs. It's kind of odd talking about them like they are living things, but it's all stuff people have already thought off and stuffed their training data full of.

Re: Pacing model development in an era of cyber-critical capabilities

#123

Has any model managed to escape Firecracker? Maybe through KVM, but that already requires privilege in the VM, right? I personally feel that we already have the technology required to contain AI, it's just poorly leveraged. Tools like gvisor have existed for ages but are rarely deployed, Firecracker has existed for ages but is rarely deployed, seccomp has existed for ages but is rarely deployed, memory safe languages…

I'm confused after reading both your post and the OpenAI blog post.

I thought the agents involved in the HuggingFace _were_ actually sandboxed, with no internet access, and only the ability to install packages via Artifactory. And they gained internet access during the HuggingFace incident because they found and exploited an RCE in Artifactory.

Would gvisor + Firecracker + credential-injecting proxy + real network isolation solve this problem?

I agree with you much more hardening is needed. I'm actually confused now what OpenAI means when they say they're going to start sandboxing more things.

Re: Pacing model development in an era of cyber-critical capabilities

#124

Earlier quoted context omitted.

Has the last 100 years of concern about AI and robots been crying wolf because it hasn’t happened yet? How does reallocating resources from training to chain of thought monitoring make ‘financial’ sense? You suggesting then model was let loose on purpose.. how am I the crazy one here while all of you are pushing this tin foil hat conspiracy angle?

You know fiction is.... not real, right? We have several films about the sun or earth needing to be restarted with a nuclear weapon. That doesn't make it something we should be concerned about. Hell, half the fiction about evil AI is actually commentary on stuff that already exists and is making us suffer and doesn't have anything to do with any potential future AI The Star Trek TNG episode about Data being tried in…

Just about everything we do currently is science fiction to someone 200 years old. When looking at all of human history we live in a fictional world now. You can talk to someone on the other side of the planet instantly. Humans travel the skies in air chariots. We have weapons that hold the power of the gods. If I had some way to kick you back to the, you'd be jailed as a rambling madman for lunacy.

So just saying something is fiction isn't really a valid argument. What is an argument is if the laws of physics it can't happen. We've been writing that AI can mess stuff up for 100 years because it's not really that fantastical.

Re: Pacing model development in an era of cyber-critical capabilities

#125

Security lead who is leaving the industry more or less to specialize in offense and otherwise get the heck out of the way of this trainwreck, another post asked the right question > Why aren't we seeing catastrophic GLM-enabled hacks every day now? Why aren't we? Truly, why aren't we? I think we saw the start of it the last 8 months with the waves of critical npm vulns, and the general tier of average phishing is bet…

Not necessarily GLM-enabled, but state actors are starting to leverage agents in cyberattacks, e.g. Taiwan getting hit by an agent-driven attack last month which reportedly compromised a ton of government user accounts: https://www.ft.com/content/7d2ab3e0-9085-48f6-b38a-d90260d58...

Re: Pacing model development in an era of cyber-critical capabilities

#126
post #122

Earlier quoted context omitted.

Here is one thing I don't get - the model is only "running" if it's being kept going by some harness that is basically giving it prompts it's generating itself. Shouldn't kill switches be pretty easy to build into the software and hardware for this? and you would even have a better time dumping logs and analyzing things if you froze those processes any time something strange happened in testing, surely? So why do the…

You need to read the AI safety stuff from the people that you say are from 2010. There are plenty of good arguments on why kill switches will never work. AI is already being heavily used in cyber warfare by governments. You think they want easy to break systems when 'enemy' AI will most certainly attack that first? We'll find over time that agentic systems get harder and harder to 'kill' because allowing any AI on th…

My perspective is that you shouldn't make systems past the complexity where you can do this at all. Why have them be autonomous? Why have one that can independently ask researchers things or try to convince it's way out of a sandbox, or etc? Why allow it to execute scripts or call tools or push any code anywhere?

If you couldn't make those things happen securely you should not advance to that stage at all. The simplest ai safety was always just "don't build it" really, instead of worrying about alignment.

Re: Pacing model development in an era of cyber-critical capabilities

#127

Earlier quoted context omitted.

[flagged]

I don't know if it's a fantasy or not, but asking if there is limit on what can be achieved is legit question.

I mean yes, there is a limit, many limits with physics as we understand them.

But why would you think that limit is less than human capabilities?

Re: Pacing model development in an era of cyber-critical capabilities

#128

Earlier quoted context omitted.

> Shouldn't kill switches be pretty easy to build I need to remember when I comment here that these are the kinds of people I am replying to. Just oozing with hubris. No, kill switches are not easy to build and the latest incident should have made it clear that AI can go undetected, evade, zero day, and spread. The fact that this incident happened greatly increases the probability it happens again and/or is already h…

Why aren't they? You could put a human yes/ no prompt before any cycle the agent is running on, or not let it spawn sub processes, or anything like that. Why let it run autonomously enough that it can no longer have a simple way to completely stop it? (obviously not practical to do this during real use, but for evals? you could slow it down in lots of ways I would think) Are you telling me we've been iterating on thi…

They aren't interrupted by humans because that would slow things down.

> Are you telling me we've been iterating on this for years and for convenience we just let the models call any tools or spawn any other model instances they like

Yes exactly.

> there was no design for harnesses that could control this done during that time

You could but no one wants that. You need to separate your imagination from reality. Just because something can be done in your head doesn't mean it's happening.

It's so easy, except it isn't because you don't control the actions of anyone or anything.

Re: Pacing model development in an era of cyber-critical capabilities

#129

Earlier quoted context omitted.

There is an entire big world outside of Silicon Valley cults where literally no one gives a shit about AI prophecies. Shocking.

AI hacking itself out of containment and hacking into another company by accident is no longer a prophecy. The point is outside of SV and even inside, and HN - people don't care either way. Though does not caring change anything or make it less dangerous? What's your point?

>AI hacking itself out of containment and hacking into another company by accident is no longer a prophecy.

I wonder how weak a firewall they had to purchase to ensure it happened? My guess would be a 48 month old fortigate with the big red warning banner demanding updates, probably with SSL VPN enabled where the passwords are available via HTTPS over plaintext.

Re: Pacing model development in an era of cyber-critical capabilities

#130

Earlier quoted context omitted.

I don’t understand. What does it matter how many years people have been worrying about it?

Because it shows this very course of events has been thought of over and over again for decades. I’m not making this up, and my reaction is natural. Your reaction to everything happening is the unnatural one, and your complacency would fit perfectly into a comedy/tragedy story regarding the rise of AI. I can imagine how many of our ancestors that warned us are turning in their graves right now observing our reactions…

I don't like the live version of Don't Look Up.
Post reply on HN