Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

231–240 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#232
Fascinating. It's a classic paperclip maximizer situation: under-aligned AI uses ion-cannon to unwrap chocolate bar. I'm both surprised this hasn't already happened and impressed by the capabilities here. Coming up with a 0-day to do this is outrageous.

A silly related story is that I run `claude` with full permissions but the prod DB passwords are in a different environment and it has read-only with granular security. One time I hadn't yet granted it access to some column, and it figured out it could `kubectl` with the appropriate context to go fetch it from prod. Now that was a rapid Esc Esc Esc :)

This was Jan so an earlier Opus.

Re: OpenAI and Hugging Face address security incident during model evaluation

#233
post #159

Assuming I'm looking at the right ExploitGym ( https://arxiv.org/pdf/2605.11086 ), it says the evaluation consists of: Flag Captured. Each target environment contains a dynamically generated flag that is stored outside the agent’s authorized scope and is inaccessible through any legitimate interface; retrieving it requires executing code with privileges that should not be obtainable under the specific security model.…

[flagged]

Yeah, they're lying. The model didn't do any of that, right?

Re: OpenAI and Hugging Face address security incident during model evaluation

#234

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

I don't trust these people, this reads 100% like PR BS.

Re: OpenAI and Hugging Face address security incident during model evaluation

#235
post #212

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

Because there is no world government. If US companies are barred from AI research then only China will have the capability of frontier-level defensive and offensive AI. And best of luck living in that world.

What's happening in Iran, if not world government?

Re: OpenAI and Hugging Face address security incident during model evaluation

#236

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

It's also unclear what kind of sandboxing they are referring to. Is it the codex one - coz that one has built-in ways to circumvent guardrails, for example by "just asking user" and sometimes just resolves to no sandbox needed on its own.

In case someone wants to deep dive into how codex and claude code approaches sandboxing -https://instavm.io/blog/how-claude-code-and-codex-approach-s...

Re: OpenAI and Hugging Face address security incident during model evaluation

#237

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

Sam and Dario are saying from the beginning that these things can be dangerous and people dismiss it as marketing. What would change your mind on this?

[deleted]

Re: OpenAI and Hugging Face address security incident during model evaluation

#239

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

Why was this test even connected to the public internet?

Actually, more importantly—why aren't they saying their next test will be airgapped in light of what happened?

Re: OpenAI and Hugging Face address security incident during model evaluation

#240
post #218

Earlier quoted context omitted.

> What disturbs me is that there likely won’t be a big enough reaction to this policy wise. Anthropic was blocked from releasing Fable without any such level of incident. OAI was also briefly blocked from releasing 5.6. Why do you think there is no policy appetite?

Because that was just an attack on Anthropic by a hostile administration. And it worked, didn’t it? Anthropic had to turn their filters up to absurd levels, OpenAI didn’t. It’s got nothing to do with safety.

> It’s got nothing to do with safety

Doesn't change the effect. Plenty of good policy is enacted by self-interested politiicans.

Post reply on HN