OpenAI and Hugging Face address security incident during model evaluation
231–240 of 1001 posts
Re: OpenAI and Hugging Face address security incident during model evaluation
#232A silly related story is that I run `claude` with full permissions but the prod DB passwords are in a different environment and it has read-only with granular security. One time I hadn't yet granted it access to some column, and it figured out it could `kubectl` with the appropriate context to go fetch it from prod. Now that was a rapid Esc Esc Esc :)
This was Jan so an earlier Opus.
Re: OpenAI and Hugging Face address security incident during model evaluation
#233Assuming I'm looking at the right ExploitGym ( https://arxiv.org/pdf/2605.11086 ), it says the evaluation consists of: Flag Captured. Each target environment contains a dynamically generated flag that is stored outside the agent’s authorized scope and is inaccessible through any legitimate interface; retrieving it requires executing code with privileges that should not be obtainable under the specific security model.…
[flagged]
Re: OpenAI and Hugging Face address security incident during model evaluation
#234I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…
Re: OpenAI and Hugging Face address security incident during model evaluation
#235I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…
Because there is no world government. If US companies are barred from AI research then only China will have the capability of frontier-level defensive and offensive AI. And best of luck living in that world.
Re: OpenAI and Hugging Face address security incident during model evaluation
#236I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…
In case someone wants to deep dive into how codex and claude code approaches sandboxing -https://instavm.io/blog/how-claude-code-and-codex-approach-s...
Re: OpenAI and Hugging Face address security incident during model evaluation
#237I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…
Sam and Dario are saying from the beginning that these things can be dangerous and people dismiss it as marketing. What would change your mind on this?
Re: OpenAI and Hugging Face address security incident during model evaluation
#238Re: OpenAI and Hugging Face address security incident during model evaluation
#239I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…
Actually, more importantly—why aren't they saying their next test will be airgapped in light of what happened?
Re: OpenAI and Hugging Face address security incident during model evaluation
#240Earlier quoted context omitted.
> What disturbs me is that there likely won’t be a big enough reaction to this policy wise. Anthropic was blocked from releasing Fable without any such level of incident. OAI was also briefly blocked from releasing 5.6. Why do you think there is no policy appetite?
Because that was just an attack on Anthropic by a hostile administration. And it worked, didn’t it? Anthropic had to turn their filters up to absurd levels, OpenAI didn’t. It’s got nothing to do with safety.
Doesn't change the effect. Plenty of good policy is enacted by self-interested politiicans.