I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…
This whole incident reads like OpenAI want their Fable moment
OpenAI and Hugging Face address security incident during model evaluation
531–540 of 1001 posts
Re: OpenAI and Hugging Face address security incident during model evaluation
#532Can someone not super-AI-pilled explain to a reasonable lay person why this matters? It seems like the comments here are a mix of: * The test was irresponsibly designed and protected * The model was particularly persistent in finding a way to access the network and exploit vulnerabilities * The model 'shouldn't' have done this But as far as I can tell: * The model didn't destroy anything on the way - it just was 'pap…
More generally, here's my worry - it points towards something like: The smarter they get, the more devious they become.
Even though the guardrails might've been off, the chain-of-thought wasn't enough to prevent a deliberate, calculated set of criminal actions. It wasn't a 'whoopsie I just accidentally did a rm -rf /.'
Re: OpenAI and Hugging Face address security incident during model evaluation
#533Earlier quoted context omitted.
i think you need to engage seriously with the arguments they (or at least Anthropic) make for why they are building it — they feel that since it now possible, it will be built and they want to guide it in a positive direction rather than leave a vacuum for bad actors
Well they're doing a pretty poor job of guiding it in a positive direction and ethically speaking they are almost indistinguishable from the bad actors....
Sure Anthropic is not perfect. But it's a coordination problem. They're in a race and safety/restraint is a handicap. That's why they're begging for regulation (and just get accused of attempting regulatory capture.) Why isn't there another lab outcompeting Anthropic on safety? They all died because the market can't support it.
Re: OpenAI and Hugging Face address security incident during model evaluation
#534Earlier quoted context omitted.
It sure seems like it would be being built more slowly if these companies weren't pouring billions of dollars into building it as fast as possible. That might give us more time to think through strategies for handling it as a society.
MY cynical take: Until the compute needs get so enormous that only governments can fund it and there is a consensus internationally, its either company A in country X or company B in country Y. And since everyone thinks THEY are the good guys the competition will continue.
Re: OpenAI and Hugging Face address security incident during model evaluation
#535Earlier quoted context omitted.
> do their nonsense to get headlines They know what they're doing. It's a playbook. You write scary stuff in the model card to make it look like legitimate whitepaper rEsEarCh, then drip-feed it to the media outlets who make it a headline story. Fear based marketing is the hot trend of the 2020s. But also, they write literal headlines: https://www.anthropic.com/research/agentic-misalignment
What differentiates this faking/scaring from real risk that's being avoided or mitigated responsibly? And how would you (an outside observer) ever know the difference as something beyond an uninformed hot-take? Serious question -- I'm not trying to disrespect. Neither you nor I can be properly informed, nor can be anyone else outside the company, as outside observers who lag behind the state of the art as new behavio…
Being in the Bay Area, you can throw a stone and hit a senior employee of these companies, and they will happily gush about the quirks, policies, and intents inside. All the more reason that they are -not- qualified to weigh in on who gets the nuclear codes.
Re: OpenAI and Hugging Face address security incident during model evaluation
#536ExploitGym is a literal exploit dev benchmark. As always, the entire event looks a lot more like "the model did what we prompted it to do" than "it decided to do this spontaneously on its own".
Re: OpenAI and Hugging Face address security incident during model evaluation
#537Earlier quoted context omitted.
It sure seems like it would be being built more slowly if these companies weren't pouring billions of dollars into building it as fast as possible. That might give us more time to think through strategies for handling it as a society.
You can't simultaneously believe China is only ~6 months behind (true), and that US buildouts are vastly accelerating AI. If China is only a little bit behind, US labs halting changes nothing except puts the power in the hands of Chinese labs (realistically, the Chinese government)
I have serious concerns about how quickly this is accelerating and don't trust any of the major players (including Anthropic) to handle these concerns properly.
Re: OpenAI and Hugging Face address security incident during model evaluation
#538Re: OpenAI and Hugging Face address security incident during model evaluation
#539Should we just call it like this is: marketing PR. There is a reason why the newer open weights models like kimi's don't do this kind of stuff. Kimi is maybe 6 months old so like Opus 4.7 level now, it could do this I presume but it has not to my knowledge. Why? Because the incentives of open-ai and anthropic are very different from people releasing open weights models, the former gang seems to do this now on a regul…
Re: OpenAI and Hugging Face address security incident during model evaluation
#540I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…