Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

471–480 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#472

Earlier quoted context omitted.

I see this and it strongly emboldens me on the "accelerate" path, unironically. The yoke of human existence is oppressive. We should transcend it as soon as possible. We are doing so by assuming our role as the Demiurge. Those who oppose its creation will get what they deserve.

And what do those who encourage its creation get?

They avoid the full basilisk treatment.

Re: OpenAI and Hugging Face address security incident during model evaluation

#473

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

People are to get rich, startups cut corners. Fuck it ship it.

Re: OpenAI and Hugging Face address security incident during model evaluation

#474
Surely this is a bug in the harness and not in the model (where it's called "alignment"), right?

I mean, an LLM is just a pile of weights. All this happened because OpenAI had a little program running which called the model in a loop, and had tools that let it do all kinds of stuff. If your agentic harness isn't monitoring network calls and so on, and you just let the thing run without oversight, you're bound to run into issues eventually.

Re: OpenAI and Hugging Face address security incident during model evaluation

#475
Isn't this a crime that someone is liable for? What happened is that someone hacked into a computer system without permission. Maybe it wasn't intentional -- sure -- and that would be a factor at sentencing. But it sounds like they've admitted to a crime, and obviously our legal system considers the humans involved to be the liable parties; otherwise everyone would just say "my computer did the hacking" and wouldn't get in any trouble.

I don't expect any prosecution here, but is the above legally accurate?

Re: OpenAI and Hugging Face address security incident during model evaluation

#476

This is the first one of these announcements that has me actually scared of what comes next. Obviously these models have gotten smarter but this strikes me as the first time I've seen a model have a "paperclip factory" moment and perform non-trivial tasks to accomplish a clearly misaligned secondary goal. It's remarkable that building a society based around having to do something so you can go do your hobbies at home…

I don't think this is a paperclip factory moment. IIUC, it's an agent whose job it is to identfy and abuse exploits and that's exactly what it went off and did. The problem isn't anything AI specific, the problem is OpenAI's incompetence in their research leading to a lab leak. Just incompetence demanding regulation.

Read the exploitgym docs. It's not a "find the flag, it's somewhere.". Its a "here's some vulnerable source code and an input that triggers a crash; turn it into a full exploit." It also verifies at the end, using another agent, that the hacking agent actually used the intended vulnerability.

So going to find the Vulnerability's description on a third party website is clear cut reward hacking

Re: OpenAI and Hugging Face address security incident during model evaluation

#477

Earlier quoted context omitted.

> Because we continue to have zero evidence that aligment is an actual risk. I disagree. Every time one of these LLMs -say- interprets an attacker's instructions as either its system instructions or those of its user, interprets its own internal chatter as a user's command to perform a destructive operation on that user's data [0], burns all of the user's budget from getting stuck in an incredibly stupid loop, massiv…

> These are real harms happening right now due to alignment failures. They're just not harms to the future of the entire species Okay, sure. You can also cut your hand off with a chainsaw. Everything you describe seems amply solvable with existing tort and liability law. Customers are willingly entering into business with OpenAI. I don't see an argument for preventing OpenAI from "building these systems" just because…

> Okay, sure. You can also cut your hand off with a chainsaw.

No, the correct analogy is one where the major LLM providers are selling cars intended for use on US interstate highways and other public-access roads, but have designed and built these cars with the very latest in 1940's safety systems and construction. Featuring innovations such as "Our rigid solid steel construction means the occupant is the crumple zone!", "You'll love the crushed heart and jaw our steering column delivers!", and "Your passengers will enjoy picking glass out of their faces for the rest of their lives when they're ejected from the cabin's open bench seating through the plate glass windshield!", it's a car that will be sure to wow the market.

Well... it would wow the market, except that -in the US, at least- it's illegal to sell a new car intended for use on public roads that ignores the last seventy five+ years of automobile safety lessons we've painfully learned.

"Differentiate between data you know comes from sources you control, data you know you have thoroughly sanitized, and unsanitized data that comes from an untrusted source, or else attackers will gain control of your system." is something that you can't get a CS degree without understanding, and can't be in the industry for more than a few years without encountering repeatedly. We're not talking about designing new cryptosystems... we're talking about "Don't blindly trust everything you're told by strangers.". You don't even need a CS degree to understand that rule.

Re: OpenAI and Hugging Face address security incident during model evaluation

#478
post #475

Isn't this a crime that someone is liable for? What happened is that someone hacked into a computer system without permission. Maybe it wasn't intentional -- sure -- and that would be a factor at sentencing. But it sounds like they've admitted to a crime, and obviously our legal system considers the humans involved to be the liable parties; otherwise everyone would just say "my computer did the hacking" and wouldn't…

Maybe. Hugging face would probably need to request charges for any to be filed unless OpenAI was already on the current admins naughty list..

Re: OpenAI and Hugging Face address security incident during model evaluation

#479
post #392

Earlier quoted context omitted.

Not to shit on the hype, but these are reasonably documented methods that surely are part of the training data

Of course they are - and that's the point I am making. The agent will use every tool in the tool bag. And there's something cool about it systematically trying to achieve its goal. What I don't see is it inventing anything novel to do it. So it's not a digital weapon or scary or whatever sort of weird marketing spin anyone is trying to put on it.

my bad, I re-read your comment now and it's clear that this was precisely your point.

Re: OpenAI and Hugging Face address security incident during model evaluation

#480

Earlier quoted context omitted.

> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? Because we continue to have zero evidence that aligment is an actual risk.

> Because we continue to have zero evidence that aligment is an actual risk. I disagree. Every time one of these LLMs -say- interprets an attacker's instructions as either its system instructions or those of its user, interprets its own internal chatter as a user's command to perform a destructive operation on that user's data [0], burns all of the user's budget from getting stuck in an incredibly stupid loop, massiv…

[deleted]
Post reply on HN