Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

501–510 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#501

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

What disturbs me is that there likely won’t be a big enough reaction to this policy wise. There’s been a relatively big reaction to Kimi K3 and Chinese open weights models, but only for financial reasons. Powerful people care about something that might pop the massive valuations of the AI companies, but not about the damage that AIs could do. Nor even about the damage that the Chinese models could do in the wrong han…

Follow the money, eg. investors and their connections to the Govt and media.

Re: OpenAI and Hugging Face address security incident during model evaluation

#502

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

Why was this test even connected to the public internet? Actually, more importantly—why aren't they saying their next test will be airgapped in light of what happened?

[deleted]

Re: OpenAI and Hugging Face address security incident during model evaluation

#503
Should we just call it like this is: marketing PR. There is a reason why the newer open weights models like kimi's don't do this kind of stuff. Kimi is maybe 6 months old so like Opus 4.7 level now, it could do this I presume but it has not to my knowledge. Why? Because the incentives of open-ai and anthropic are very different from people releasing open weights models, the former gang seems to do this now on a regular basis.

There are a few things that perplex me even more:

1. If you are going to eventually publicly release models that are trained to behave according your spec or AI-Constitution to maintain coherent behavior**, why on earth would you want to tell anyone it can do this?

2. Do they have another GPT 5.6 trained to not obey a different constitution/spec to do this kind of hacking? Because that makes no sense since you would never release it.

3. And if this is a constitution obeying model, I am also curious what they did to it to get it to do this hack without serious pushback from the model's training. Whenever I have tried to get codex/claude to do a vulnerability scan of my own servers it always refuses constantly.

** I know spec based training has its limitations, but its all we have and atleast one knows what the model's persona is and what its value system is. But there is no reason you would make one model do that while letting another one be a crazy hacker. Its well known if you fine tune a model to change one part of its personal other often unrelated parts of it suffer from safety issues.

Re: OpenAI and Hugging Face address security incident during model evaluation

#504
post #353

Earlier quoted context omitted.

Headline? It was buried in a model card. They just honestly report not-quite-incident because it's quite close to the incident OpenAI had. Nothing wrong with it.

> do their nonsense to get headlines They know what they're doing. It's a playbook. You write scary stuff in the model card to make it look like legitimate whitepaper rEsEarCh, then drip-feed it to the media outlets who make it a headline story. Fear based marketing is the hot trend of the 2020s. But also, they write literal headlines: https://www.anthropic.com/research/agentic-misalignment

What differentiates this faking/scaring from real risk that's being avoided or mitigated responsibly? And how would you (an outside observer) ever know the difference as something beyond an uninformed hot-take?

Serious question -- I'm not trying to disrespect. Neither you nor I can be properly informed, nor can be anyone else outside the company, as outside observers who lag behind the state of the art as new behaviors emerge, right?

Re: OpenAI and Hugging Face address security incident during model evaluation

#505

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

What disturbs me is that there likely won’t be a big enough reaction to this policy wise. There’s been a relatively big reaction to Kimi K3 and Chinese open weights models, but only for financial reasons. Powerful people care about something that might pop the massive valuations of the AI companies, but not about the damage that AIs could do. Nor even about the damage that the Chinese models could do in the wrong han…

> There’s been a relatively big reaction to Kimi K3 and Chinese open weights models, but only for financial reasons.

Let's be honest: it's financial and national security reasons.

China has a long and storied history of hacking attacks on American and western targets.

There are other parts of the world that make open weight models; Mistral is a European option. You don't see the worry about that because most people in the US are used to existing in a world order where European powers are considered ambivalent to the US at worst and holders of a special political relationship at best.

If Mistral had the same backing that Chinese AI companies did, there probably wouldn't be as much hemming and hawing. Sure, American companies would take a haircut, but that haircut wouldn't be seen as a move towards software hegemony built on top of manufacturing hegemony. It'd just be you calling into Paris or Frankfurt to talk to your vendor in the future.

Re: OpenAI and Hugging Face address security incident during model evaluation

#506

Earlier quoted context omitted.

> all that regulation will do at this point is help the incumbents who are failing This depends on the specific regulation. The datacentre moratoria probably give open-weight models time to catch up by tempering the extent to which the leading companies can turn their capital advantage into market share.

> datacentre moratoria What infrastructure will these open weight models be trained on?

Meituan’s 1.6T LongCat was trained entirely on Huawei training cards.

DeepSeek, GLM, Qwen and others are also actively working on similar replacement.

Re: OpenAI and Hugging Face address security incident during model evaluation

#507
post #154

Earlier quoted context omitted.

Oh... if Sam and Dario say so, then it must be true.

About their creation? Yes as most of inventors about their invention usually

Altman is an enabler, not an inventor

Re: OpenAI and Hugging Face address security incident during model evaluation

#508

Earlier quoted context omitted.

Lol this has to be a troll, I've never seen something so wildly, obviously, incredibly wrong. You can debate all you want if alignment is possible . That is a valid discussion. But it's trivial to demonstrate that alignment is a problem .

> can debate all you want if alignment is possible. That is a valid discussion. But it's trivial to demonstrate that alignment is a problem ...how is an impossible thing supposed to be a problem?

Alignment is something you want, so if you're not confident that it's possible, that sure sounds like a problem

Re: OpenAI and Hugging Face address security incident during model evaluation

#509

Earlier quoted context omitted.

i think you need to engage seriously with the arguments they (or at least Anthropic) make for why they are building it — they feel that since it now possible, it will be built and they want to guide it in a positive direction rather than leave a vacuum for bad actors

If this is how the 'good guys' act I think I'd rather take my chances with the bad actors...

It all feels a bit like Dr. Strangelove just the Ai arms race version. Crazies all around.

Re: OpenAI and Hugging Face address security incident during model evaluation

#510
post #218

Earlier quoted context omitted.

> What disturbs me is that there likely won’t be a big enough reaction to this policy wise. Anthropic was blocked from releasing Fable without any such level of incident. OAI was also briefly blocked from releasing 5.6. Why do you think there is no policy appetite?

> Why do you think there is no policy appetite? Because China seems pretty eager to serve the rest of the world's needs if the USA doesn't stop their idiotic "safety" nonsense.

They won’t when citizens run AI on their phone that contradicts Xi thought
Post reply on HN