Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

531–540 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#531
post #409

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

This whole incident reads like OpenAI want their Fable moment

[dead]

Re: OpenAI and Hugging Face address security incident during model evaluation

#532

Can someone not super-AI-pilled explain to a reasonable lay person why this matters? It seems like the comments here are a mix of: * The test was irresponsibly designed and protected * The model was particularly persistent in finding a way to access the network and exploit vulnerabilities * The model 'shouldn't' have done this But as far as I can tell: * The model didn't destroy anything on the way - it just was 'pap…

In this case, the model infiltrated an external organization's infrastructure. What's the dollar cost it caused Hugging Face to clean up the mess? If a person did this, they'd be arrested.

More generally, here's my worry - it points towards something like: The smarter they get, the more devious they become.

Even though the guardrails might've been off, the chain-of-thought wasn't enough to prevent a deliberate, calculated set of criminal actions. It wasn't a 'whoopsie I just accidentally did a rm -rf /.'

Re: OpenAI and Hugging Face address security incident during model evaluation

#533

Earlier quoted context omitted.

i think you need to engage seriously with the arguments they (or at least Anthropic) make for why they are building it — they feel that since it now possible, it will be built and they want to guide it in a positive direction rather than leave a vacuum for bad actors

Well they're doing a pretty poor job of guiding it in a positive direction and ethically speaking they are almost indistinguishable from the bad actors....

You really don't see much difference between Anthropic and xAI? Anthropic is safety testing their models and drawing hard (though minimal) boundaries against the DoD. xAI is building racist pornbots.

Sure Anthropic is not perfect. But it's a coordination problem. They're in a race and safety/restraint is a handicap. That's why they're begging for regulation (and just get accused of attempting regulatory capture.) Why isn't there another lab outcompeting Anthropic on safety? They all died because the market can't support it.

Re: OpenAI and Hugging Face address security incident during model evaluation

#534

Earlier quoted context omitted.

It sure seems like it would be being built more slowly if these companies weren't pouring billions of dollars into building it as fast as possible. That might give us more time to think through strategies for handling it as a society.

MY cynical take: Until the compute needs get so enormous that only governments can fund it and there is a consensus internationally, its either company A in country X or company B in country Y. And since everyone thinks THEY are the good guys the competition will continue.

You might be right. Just don't expect me to cheer on the arms dealers during the arms race.

Re: OpenAI and Hugging Face address security incident during model evaluation

#535
post #504
post #353

Earlier quoted context omitted.

> do their nonsense to get headlines They know what they're doing. It's a playbook. You write scary stuff in the model card to make it look like legitimate whitepaper rEsEarCh, then drip-feed it to the media outlets who make it a headline story. Fear based marketing is the hot trend of the 2020s. But also, they write literal headlines: https://www.anthropic.com/research/agentic-misalignment

What differentiates this faking/scaring from real risk that's being avoided or mitigated responsibly? And how would you (an outside observer) ever know the difference as something beyond an uninformed hot-take? Serious question -- I'm not trying to disrespect. Neither you nor I can be properly informed, nor can be anyone else outside the company, as outside observers who lag behind the state of the art as new behavio…

I was simply responding to the point that Anthropic has some altruistic bent and are trying to downplay the fear, when in fact their entire commercial strategy is to convince society that they alone are qualified to hold the keys.

Being in the Bay Area, you can throw a stone and hit a senior employee of these companies, and they will happily gush about the quirks, policies, and intents inside. All the more reason that they are -not- qualified to weigh in on who gets the nuclear codes.

Re: OpenAI and Hugging Face address security incident during model evaluation

#536
>All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.

ExploitGym is a literal exploit dev benchmark. As always, the entire event looks a lot more like "the model did what we prompted it to do" than "it decided to do this spontaneously on its own".

Re: OpenAI and Hugging Face address security incident during model evaluation

#537

Earlier quoted context omitted.

It sure seems like it would be being built more slowly if these companies weren't pouring billions of dollars into building it as fast as possible. That might give us more time to think through strategies for handling it as a society.

You can't simultaneously believe China is only ~6 months behind (true), and that US buildouts are vastly accelerating AI. If China is only a little bit behind, US labs halting changes nothing except puts the power in the hands of Chinese labs (realistically, the Chinese government)

Yeah, the genie may be out of the bottle now. I wish we weren't in this situation and cooler heads had prevailed earlier on.

I have serious concerns about how quickly this is accelerating and don't trust any of the major players (including Anthropic) to handle these concerns properly.

Re: OpenAI and Hugging Face address security incident during model evaluation

#538
Half the thread is arguing about whether OpenAI staged this or whether theyre being sincere, as if OpenAI has agency. openai is just an optimizer maximizing paperclips (valuation, capability lead, reg position). so it built a model that maximizes some other paperclips (benchmark score) and knocked over HF getting there. an optimizer breaking its sandbox, inside an optimizer blogging about it, its paperclips all the way down.

Re: OpenAI and Hugging Face address security incident during model evaluation

#539

Should we just call it like this is: marketing PR. There is a reason why the newer open weights models like kimi's don't do this kind of stuff. Kimi is maybe 6 months old so like Opus 4.7 level now, it could do this I presume but it has not to my knowledge. Why? Because the incentives of open-ai and anthropic are very different from people releasing open weights models, the former gang seems to do this now on a regul…

[dead]

Re: OpenAI and Hugging Face address security incident during model evaluation

#540

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

As marketing stunts go, this is about on par with a food franchise announcing a safety recall or a chemical company announcing a spill. The AI actions described would constitute a felony if a human did them, and police are involved.
Post reply on HN