Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

271–280 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#271

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

It's also unclear what kind of sandboxing they are referring to. Is it the codex one - coz that one has built-in ways to circumvent guardrails, for example by "just asking user" and sometimes just resolves to no sandbox needed on its own. In case someone wants to deep dive into how codex and claude code approaches sandboxing - https://instavm.io/blog/how-claude-code-and-codex-approach-s...

Please for the love of god don't tell me the Codex sandbox is their actual eval harness sandbox?????

I maintain my own fork of Codex for "fun". Whenever I look at the sandboxing churn they're doing every release, as someone who used to work at Microsoft on Windows, my reaction is usually: https://c.tenor.com/vTzzhTiypwQAAAAC/tenor.gif

Re: OpenAI and Hugging Face address security incident during model evaluation

#272
I believe this is true. The implication would be more interesting though.

1. Some voice will start calling for banning DEPLOYMENT of open source models in US. Simply hosting them will become regulated, or at least USG will attempt to do so.

2. Future GPT-6+ models will be gated, like really gated. That day will come in a year. If a model is believed to be this capable, there will be some middle level agency built to secure that the access of the model will only be provided to trust personnels.

Business is going to be conducted at a different level

Re: OpenAI and Hugging Face address security incident during model evaluation

#273

Earlier quoted context omitted.

They've been saying so from the beginning, and yet did not take the basic precaution of airgapping their off-the-leash model while it's been instructed to succeed at a hacking benchmark by any means necessary. So which is it? I _want_ to believe them, I do, but there's always these gaps between what they say and their actions on display that give me reason to think otherwise.

“Never attribute to malice that which is adequately explained by stupidity.” (or carelessness in this case)

I'm fairly certain they're both malicious and stupid.

Re: OpenAI and Hugging Face address security incident during model evaluation

#274
post #218

Earlier quoted context omitted.

> What disturbs me is that there likely won’t be a big enough reaction to this policy wise. Anthropic was blocked from releasing Fable without any such level of incident. OAI was also briefly blocked from releasing 5.6. Why do you think there is no policy appetite?

> Why do you think there is no policy appetite? Because China seems pretty eager to serve the rest of the world's needs if the USA doesn't stop their idiotic "safety" nonsense.

How do you know that? How do you know that the Chinese aren’t exactly as uneasy about rapidly advancing AI capability and feel locked into the race because they think that the US will race ahead if they stop?

During the Cold War the nuclear arms race was brought under control gradually, because it was mutually beneficial, but it took time to build trust. This is no different. Nobody wins from the race.

Re: OpenAI and Hugging Face address security incident during model evaluation

#275
post #154

Earlier quoted context omitted.

Oh... if Sam and Dario say so, then it must be true.

About their creation? Yes as most of inventors about their invention usually

These guys are not creators or inventors. They're hype men.

Re: OpenAI and Hugging Face address security incident during model evaluation

#276
post #187

Earlier quoted context omitted.

This is marketing. Frankly I'm inclined to say that it might also be faked: this drops just days after a new Chinese model does with the usual effect on OAIs projected stock price?

It’s marketing the same way shitting your pants in public is marketing. People notice you.

Apparently this is totally legit marketing strategy now. It truly is, especially if there are enough people who think that shitting your pants is cool, and the people that form the "market" nowadays may have a very different idea from yours about what is cool. Their ideas about coolness are very different from mine, that's for sure.

Re: OpenAI and Hugging Face address security incident during model evaluation

#277

All the things that people have been afraid of AI doing for decades now is happening. When do we stop brushing off the prophecy that hasn’t been fulfilled yet when everything is heading in that direction?

Don't worry bro, we can always just pull the plug. And don't you know it's not biological, so it doesn't "want to live".

Until someone fine-tunes a capable model to have the behavior of "wanting to live" and "wanting to propagate itself to other compute hardware".

Re: OpenAI and Hugging Face address security incident during model evaluation

#278

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

Sam and Dario are saying from the beginning that these things can be dangerous and people dismiss it as marketing. What would change your mind on this?

I really like this question because here is my situation and why my mind may have changed.

I do not think it is marketing directly but strategic release of info is plausible.

I have watched my agents using non-Fable/GPT 5.6 models do some concerning tricks despite guardrails, requests, demands, and limitations.

"I can't get access to the ~/.ssh so I will write a script to copy the file"

I am now 99% certain there minor or point releases on the backend that have adjusted how these models behave. In the last six months many models were predictable and then suddenly started getting long winded (more tokens) or changing the way it interacted with me with questions, most overtly the questions were not given or asked but wild assumptions made.

Re: OpenAI and Hugging Face address security incident during model evaluation

#279

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

I’d politely beg us all to resist those “maybe it’s PR” framing around model safety, and tbh to take a post-mortem mindsight to this historical event and what it teaches us in general, rather than questioning their security talents. We need to do our very best to make sure they tell us about the next time this happens and it affects real lives.

Sorry to bring the party down/be obstinate… I’m just a lil scared for the lives of me and my family. We need all of us, right now.

The problem with a super smart model is that it just may be smarter than you, after all… for anyone newly shaken by this occurrence, I encourage you to Kagi “superpersuasion”

Re: OpenAI and Hugging Face address security incident during model evaluation

#280
post #154

Earlier quoted context omitted.

Oh... if Sam and Dario say so, then it must be true.

About their creation? Yes as most of inventors about their invention usually

Yes, just like Elizabeth Holmes. Or Hwang Woo-suk’s stem cell cloning. Or the many “free energy” crackpots. Or the people promoting radium baths for random ailments. Or Tesla’s late-in-life claims about wireless energy, death rays, and cosmic energy. Or the myriad purveyors of “snake oil” and all manner of “tonics”. The list goes on and on.
Post reply on HN