Live data from Hacker News

Timeline of the OpenAI accidental attack against Hugging Face

simonwillison.net

21–30 of 260 posts

Re: Timeline of the OpenAI accidental attack against Hugging Face

#21

Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose? If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and…

There are trade-offs here:

Give up too early -> users will get annoyed because the task would have been solvable if the model pushed harder.

Give up too late -> collateral damage while completing the task A.K.A. misalignment.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#22

Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose? If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and…

> If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and say “I’m not sure how to proceed next”

This goes against the goal of "solve this math problem that no human was able to solve for 80 years, do NOT give up, even if you know it's unsolved and really hard"

Re: Timeline of the OpenAI accidental attack against Hugging Face

#24
post #10
post #2

Isn't this a show of security negligence rather than of exceptional agent capabilities? Don't get me wrong, I am pretty impressed that an agent was able to use these vulnerabilities. But I am way more impressed by the vulnerabilities...

Yes. It is very easy to add to the instructions "for every potential exploit you discover and use, document them as you go into this repository" and have alerting there. The fact that they did not do this means they wanted to be surprised, and have plausible deniability on their side when things inevitably blow up. And for my fellow engineers who would think "oh no, they wouldn't do that". Remember that these places…

/s?

"Btw don't turn the planet into paperclips"

Re: Timeline of the OpenAI accidental attack against Hugging Face

#25
post #2

Isn't this a show of security negligence rather than of exceptional agent capabilities? Don't get me wrong, I am pretty impressed that an agent was able to use these vulnerabilities. But I am way more impressed by the vulnerabilities...

> Isn't this a show of security negligence rather than of exceptional agent capabilities?

Seems to me you could say this about all enterprise adoption of "AI" since 2023.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#26
post #2

Isn't this a show of security negligence rather than of exceptional agent capabilities? Don't get me wrong, I am pretty impressed that an agent was able to use these vulnerabilities. But I am way more impressed by the vulnerabilities...

OpenAI reported the Artifactory vulnerability, patched it, then the agents immediately found a new zero day.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#27

Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose? If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and…

They certainly want their models to be good at finding and patching vulnerabilities. Being good at hacking may be necessary in that goal, or rather, making it worse at hacking may also make it worse at defensive actions too.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#29
From the outside, it looks like OpenAI got exactly the kind of event they could market the hell out of to demonstrate the capability of the model.

But the event itself only seems possible because they failed to properly monitor and isolate the environment in the first place. To me, it looks like their job is to market the model, not take security seriously.

The model is obviously impressive, but we already knew that. I personally don’t like how the containment failure becomes part of the mythology of how capable the model is, rather than an environment engineering failure.

At the end of the day, it’s not like Hugging Face is critical infrastructure. But there need to be real consequences for stuff like this so that OpenAI is incentivized to mature as an organization and take security more seriously.

At this point, this incident is just security porn and entertainment for developers

Re: Timeline of the OpenAI accidental attack against Hugging Face

#30
post #12

Would love to see a cat and mouse game being played by openai versus anthropic, out in the open.

Military has a phrase for the outcome - collateral damage.

> Yes, I just hacked into AWS and shut down all of the data-centers, because it's where Anthropic Mythos servers are hosting the model.

Post reply on HN