Live data from Hacker News

OpenAI’s accidental attack against Hugging Face is science fiction that happened

simonwillison.net

81–90 of 475 posts

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#81
post #51
post #18

Earlier quoted context omitted.

You know this makes OpenAI look really bad , right? Hugging Face had to tell all of their users, many of them paying customers: > As a precaution, we recommend rotating any access tokens and reviewing recent activity on your account. If you believe you are affected, or want to report a security concern, contact us at security@huggingface.co. HF also said this, I'd be very interested to hear how that got resolved! > F…

I could fully see them thinking the incident disclosed yesterday would have made them look good ("wow, OpenAI's models are so capable!"). That it didn't occur to them to discuss specific preventative measures to be taken in the future (airgapping as a foolproof one already familiar to the CTF world, anyone?) indicates to me they're not taking their job seriously; they are the ones treating this as a marketing charade…

> airgapping as a foolproof one

How would the model get any packages that it thinks it needs to complete the task at hand? Not a well-specified task that those tasking it could anticipate and provide all resources up front, but one of discovery.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#82
Replace LLM mentions with actual humans and this sounds a lot more serious: Rouge employees break into another company to steal hackathon answers (pinky promise)?

That's not a marketing stunt at all, if anything, more of a call for better accountability on agentic work in general.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#83
post #53

I can’t help but feel all warm and fuzzy with my head in the sand and getting a shout out in TFA for calling it marketing. We don’t currently and probably won’t ever fully understand the conditions that precipitated these events. That and the timing of this event is going to make it look suspicious to a lot of people. The truth of how it happened doesn’t matter. The attention around this will be used to create the ki…

The CEO of a cyber security company was already making comments about how this is a watershed moment for AI security. The hype machine continues (and my stocks go up)

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#84

The title (currently "OpenAI's accidental cyberattack against Hugging Face is science fiction") suggests some information had been hidden that makes the incident less significant than claimed. The article argues the opposite, and the last two words of the full title are "that happened."

I think you were reading "... is science fiction" the wrong way out of two possible interpretations. I don't think "it's science fiction" meant "it's made up". I think it meant "it sounds like something you'd read in science fiction (except this time it's something that actually happened)".

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#85
post #18

>To those people I say pull your heads out of the sand—you’re now including Hugging Face in your conspiracy theories, just so you can deny the crescendo of evidence here! Not really. I get the impression that they shoved their cyber available models behind a really shithouse proxy and went "Oh I sure hope it doesnt exploit the proxy and escape to hack huggingface" and that doesn't require Huggingface to be a willing…

You know this makes OpenAI look really bad , right? Hugging Face had to tell all of their users, many of them paying customers: > As a precaution, we recommend rotating any access tokens and reviewing recent activity on your account. If you believe you are affected, or want to report a security concern, contact us at security@huggingface.co. HF also said this, I'd be very interested to hear how that got resolved! > F…

>You know this makes OpenAI look really bad, right?

The target audience is regulators. They want to look like the smart guys really concerned about AI safety, when they come asking for open weights models to be banned and for other regulations to cement in their moat.

They want this to look like a demon core incident. Bomb and Nuclear reactors still got built.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#86
I think points that deserve more attention in the current public discourse are:

- This should be a huge wakeup call for everybody.

- We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something.

- It also shows apparent lack of competence and oversight from OpenAI: how is it that they didn't quickly find that agent is breaking the sandbox and roaming their internal network?

- What if in the future similarly misaligned AI agent tries to export its own weights and hack and clone itself into instances at various cloud hosting providers? Suddenly we might be dealing with a persistent threat harder to contain.

- The OpenAI post about this shows surprising lack of ability to see the seriousness of all this.

- For their models this isn't just an unlucky incident: it seems there have been multiple such cases recently, e.g. https://openai.com/index/safety-alignment-long-horizon-model...

- The fact that it happened again seems to show their lack of ability to derive useful oversight measures.

- Or they just don't care enough?

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#87

Earlier quoted context omitted.

If a criminal can escape a prison, that's usually negligence on part of the prison staff. Now suppose the criminal can think 1000 times faster than a typical human, can act 1000 times faster than a typical human, and knows 1,000,000 times more than a typical human. Is the prison staff still at fault for not preventing the outbreak?

I can see my carefully-worded post is getting d*wnvoted, and I see from your comment why: it's being skimmed and people are assuming I'm talking about blame. To address your point though, if every brick in the prison were made by a different person, and the prison "architects" simply glued random bricks together, I think that's closer to what we have in software right now.

> I can see my carefully-worded post is getting dwnvoted, and I see from your comment why: it's being skimmed and people are assuming I'm talking about blame.

No, you're being "dwnvoted" as you said because you're wrong, multiple times in multiple different ways in your "carefully-worded post".

>Everyone is getting AI psychosis over this one. There really isn't that much to see here.

Implying that an AI hacking it's way out of a system and into another has nothing to do with AI. When clearly it does - it's an AI that did it.

>OpenAI disabled all of the safeguards on a model that was likely trained specifically to exploit systems, and the prompt was probably something like "you're a hacker, try to hack this",

No the goal this evaluation was not to try to hacks, it was to see if an already known hack could be turned into a useable exploit. Ie "turn these ingredients in this basket into a cake." Not "go off and grow, harvest and mill your own flour, to bake a pasta dish, to bribe some to get access to a cake someone else already baked."

> and surprise! It correctly figured out that it's a test and it did hacker things.

"Doing hacker things" completely misses the point. That's just barely more accurate than dismissing it because "it uses a computer and surprise it did computer things".

> The real story here is: Some people have been sounding the alarm for years that modern software is full of holes, and finally there's nothing left to hide behind. Pretending they don't exist is no longer sustainable.

No that's not the real story. As you said that's been the case for years, so that's not the story here.

The story here is that they built a very powerful, uncontrolled agent with strong paper-clip maximizing tendencies.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#88

Earlier quoted context omitted.

If a criminal can escape a prison, that's usually negligence on part of the prison staff. Now suppose the criminal can think 1000 times faster than a typical human, can act 1000 times faster than a typical human, and knows 1,000,000 times more than a typical human. Is the prison staff still at fault for not preventing the outbreak?

I can see my carefully-worded post is getting d*wnvoted, and I see from your comment why: it's being skimmed and people are assuming I'm talking about blame. To address your point though, if every brick in the prison were made by a different person, and the prison "architects" simply glued random bricks together, I think that's closer to what we have in software right now.

It is downvoted because you first talk about "AI psychosis", then acknowledge the very issues people are concerned about.

If our prisons are all random bricks glued together, that doesn't change the practical problem caused by latest AI models more easily exploiting this.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#89
post #49

Where does this leave formal verification? Are we just shit-out-of-luck at this point? You can formally verify everything about an airplane's code, but if any of that is wrong, ChatGPT might decide that the best way to help you win the Nobel Prize is to take down the airplane your chief rival for the prize is currently in.

I think formal verification has never looked better. The main reason formal verification has never really taken off is that it's difficult. LLMs are significantly more familiar with Lean and Rocq and TLA+ than most software engineers. I think the cost of trying to build systems that adopt formal verification may have just dropped low enough that companies will consider them when previously the ROI didn't look like it…

I have had a lot of sympathy for this statement, LLMs could lower the bar to use of formal methods. But thinking it over in the context of BDD-driven development I am no longer really that sure. Compare two scenarios: A) from a specification an AI agent develops a usual piece of code along with a BDD-style testsuite passed and B) same AI also delivers a formal test (Lean/Rocq..) and successfully executes and passes it.

Will human judgment really consider scenario B) more credible than A) ? By so much that it is worth the effort ?

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#90
post #86

I think points that deserve more attention in the current public discourse are: - This should be a huge wakeup call for everybody. - We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something. - It also shows apparent lack of competence and oversight from OpenAI: how is it that they didn't quickly find that agent is breaking the sandbox…

It's a fake PR issue. It's hardly the first time this happens, but of course OpenAI, with its IPO now more in doubt than ever, had to claim this (and, once again, I have trouble believing Sam Altman choosing this: this could lead to OpenAI getting regulated, which has at least as much potential to lower their IPO price as to raise it).

But there have been messages about LLMs, especially coding agents, "grabbing root" etc many times. I have experienced such an oops. Such a hack has happened and been reported on this very site:

https://news.ycombinator.com/item?id=48348578

Post reply on HN