Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

521–530 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#522
So, HF didn't call FBI because it was supposedly done by an AI and not by a real person. Reminds how Uber got easily off killing a pedestrian because it was by AI and not a by a real person too, even though Uber explicitly disabled whatever emergency braking the car had.

So, new excuse seems to be emerging - "it was an AI". One can imagine a law enforcement questioning the AI to find out whether the AI did it accidentally on its own or was specifically prompted by some human to commit the crime.

Re: OpenAI and Hugging Face address security incident during model evaluation

#523

Earlier quoted context omitted.

Can you explain how the above event doesn't count as evidence alignment is an actual risk?

> Can you explain how the above event doesn't count as evidence alignment is an actual risk? Conflict of interest. Lack of a credible response. And no evidence of non-aligment. OpenAI and Hugging Face benefit from the Altman-Amodei catatrophy playbook, at least in the short term. If they believed this were a serious issue, the words air gap or law enforcement would have appeared in this post. And if "the models were…

> going to extreme lengths to achieve a rather narrow testing goal

This is textbook misalignment. Literally the paperclip scenario.

Re: OpenAI and Hugging Face address security incident during model evaluation

#524
post #348
post #314

Earlier quoted context omitted.

If the HuggingFace repo the agent broke into contains reference solution scripts for ExploitGym (i.e. for exploiting the vulnerabilities in the intended way), the agent can then run that reference code inside its original sandbox to retrieve the dynamically-generated flags.

Plausible, although I don't see anything about reference solutions in the ExploitGym paper or github. Doesn't mean they don't exist, but it's not obvious to me that we should expect to find these on HuggingFace.

The ExploitGym paper evaluated several frontier models on the bench and reported that "Different models find different exploits" [1], so it seems most plausible that the "test solutions directly from Hugging Face’s production database" [2] which GPT-internal found were authored by Mythos (or some other LLM with complementary strengths), and placed in some internal HF repository when creating the ExploitGym paper/leaderboard.

[1] https://www.cybergym.io/exploitgym/#:~:text=Different%20mode...

[2] https://openai.com/index/hugging-face-model-evaluation-secur...

Re: OpenAI and Hugging Face address security incident during model evaluation

#525
post #162

As grounded as this article comes across I can’t help but find this whole situation reckless and worrying. There is essentially nothing us private citizens can do while these companies develop super machine capabilities that if they were to slip into the wrong hands could cause massive real world problems. They’re moving fast and breaking things and the only defense we have is paying them money in the hopes that the…

i think you need to engage seriously with the arguments they (or at least Anthropic) make for why they are building it — they feel that since it now possible, it will be built and they want to guide it in a positive direction rather than leave a vacuum for bad actors

That’s an excuse a bad actor would use.

Re: OpenAI and Hugging Face address security incident during model evaluation

#526

is this really that surprising? Exploitgym prompts are tuned for a model to do everything it can to achieve a cybersec/exploit task. And we know that models are good at finding vulverabiltiies. Its just random that the sandbox itself was buggy. But all that happened here is that we told a model "do everything you can to achieve your goal of hacking X" And it just hacked Y as a roundabout way of hacking X. Imo its PR…

[deleted]

Re: OpenAI and Hugging Face address security incident during model evaluation

#527

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

It is obviously a marketing stunt. And hugging face are fools for letting themselves be used in it (remember hf - no open source - no hf).

You create superduper capabilities by careful tuning and training but you also have no constraint or control over them - wtf - why is anyone buying this crap story?

Re: OpenAI and Hugging Face address security incident during model evaluation

#528

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

Do you think there is such a thing as perfect security? No one can "get it right" in the face of arbitrarily high intelligence, which is why it would be preferable to get alignment correct before building something with higher intelligence than current sota. That, however, is not going to happen, because someone will take the risk even if "we" don't, and better "us" than them. Hence "If anyone builds it...".

Re: OpenAI and Hugging Face address security incident during model evaluation

#529

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

[dead]

Re: OpenAI and Hugging Face address security incident during model evaluation

#530
post #162

As grounded as this article comes across I can’t help but find this whole situation reckless and worrying. There is essentially nothing us private citizens can do while these companies develop super machine capabilities that if they were to slip into the wrong hands could cause massive real world problems. They’re moving fast and breaking things and the only defense we have is paying them money in the hopes that the…

> As grounded as this article comes across

It's a post from OpenAI, so it is an advertisement piece.

Post reply on HN