Live data from Hacker News

OpenAI’s accidental attack against Hugging Face is science fiction that happened

simonwillison.net

151–160 of 475 posts

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#151
post #86

I think points that deserve more attention in the current public discourse are: - This should be a huge wakeup call for everybody. - We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something. - It also shows apparent lack of competence and oversight from OpenAI: how is it that they didn't quickly find that agent is breaking the sandbox…

> The fact that it happened again seems to show their lack of ability to derive useful oversight measures. I think OpenAI likes the attention and did not try particularly hard to constrain the setup, even when it went off the rails. Also, the whole point is to see how good the models are at exploiting stuff when unconstrained . Turns out: quite good, as expected. Let me restate what I said in the other thread: Would…

It is not possible one of their extraordinarily high paid engineers did not know how to deploy an airgapped environment for the models to run in. Even if somehow true, they also clearly failed to contract specialists like myself to advise them on how to airgap software properly. Models will not break the laws of physics.

They simply thought "Running in a VM/Container is easier and probably fine".

And the next 1000 escapes will be for the same reason, because negligence is quick and thus more profitable.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#152
post #144

Currently trying to avoid an open weight model ban while OpenAI, who is closed, lets theirs run wild on the internet causing harm to another company because they do not understand how to airgap things. Cool.

That's the fun part. They do know how to airgap things. The models are outsmarting already pretty smart people!

I don’t think this was airgapped according to their admission.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#153
> There will inevitably be some people who dismiss this story as a dishonest marketing trick by OpenAI to make their models sound terrifyingly effective. I found 81 instances of the term “marketing” in the Hacker News discussion of the incident. To those people I say pull your heads out of the sand—you’re now including Hugging Face in your conspiracy theories

I’m one of those people who remains sceptical. Not about whether this happened. But about what it means. Like, how impressive [EDIT: tight] was the sandbox this model was in? Did the researchers really have no clue what was happening until days ex post facto?

So yes, I think something happened. But I want independent corroboration before I act on it. That isn’t the same as putting one’s head in the sand. It’s just demanding extraordinary evidence for an extraordinary and self-serving claim being made by a serial liar. The specifics matter for whether these models are a new HEU, or if they’re closer to a dangerous (but valuable) industrial process.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#154
post #86

I think points that deserve more attention in the current public discourse are: - This should be a huge wakeup call for everybody. - We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something. - It also shows apparent lack of competence and oversight from OpenAI: how is it that they didn't quickly find that agent is breaking the sandbox…

How do you "hack a lab and synthesize something"?

It’s complete science fiction so they’re allowed to say anything. They could have said the AI will upload itself to the internet, start self replicating, hack the stock market. Whatever they want because it’s all made up nonsense.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#155
post #129

The technology held by private AI companies is warfare-capable technology. Imagine the prompt: "Use all available resources to disable the power grid of ." The resource cost that prevents scaling up such a war machine is, what, just the cost of building data centers and its ongoing power bill? Cheap and easy compared to nuclear infrastructure. Governments should immediately begin leveraging this technology on the def…

> "Use all available resources to disable the power grid of ." This is like telling a team of highly qualified spies to do the same. You can ask, but whether it will succeed depends on the competency of those who established the infrastructure under attack. Sometimes the resources spent will not yield any huge vulnerabilities. > Governments should immediately begin leveraging this technology on the defense side (lite…

> Over-regulating AI wouldn't be the equivalent of limiting nuclear arms. With your analogy, which I don't think is the best one to make, it would be like regulating the study of nuclear physics.

We do, I believe, regulate uranium enrichment (the equivalent on building larger SOTA models), so while you are perfectly free to study theoretical physics and even run very large collider experiments (the equivalent of improving RLHF with DPO), it is generally frowned upon to go full Edward Teller and advocate for scientific experiments requiring detonating thermonuclear weapons in hurricane clouds (the equivalent of, well, see OP).

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#156
The thing with me and this is that the teams that were competing in the DARPA Grand Cyber Competition all had this capability, like, last year.

All the attention has been on software security, of extracting the next marginal vulnerability out of heavily-scrutinized large codebases. In the actual professional field of infosec, that's a speciality; another specialty is network pentests and red-teaming, which exploits misconfigurations and seeks out weakest-link software (rather than exhaustively fishing for the next kernel LPE or whatever).

Red teaming and netpen work is probably substantially easier for models than software security; it costs less context, but is also much more explicitly an implicit search problem where win conditions are just spotting stupid stuff that humans missed.

My visceral reaction to this is that with the right harness, you probably could have replicated this with an open-weights model last year. (I'm saying this as confidently as I am because a CGC team leader agreed with me about it yesterday).

I think people forget that the harness work we're considering here --- I don't know anything about OpenAI's harness or ExploitGym or whatever --- are basically not new; people have been developing automated exploitation and pivoting toolkits for decades, and scanners long before that. So the idea of a tool getting 0.0.0.0/0 as a target list instead of 192.168.1.0/24 and then busting up a bunch of random people's computers: not really very startling.

Obviously, LLMs give those kinds of scanners an intentionality they wouldn't have had before. But as a person who keeps a computer science perspective on security stuff, I don't know that it gives them capabilities they didn't have.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#158
post #129

The technology held by private AI companies is warfare-capable technology. Imagine the prompt: "Use all available resources to disable the power grid of ." The resource cost that prevents scaling up such a war machine is, what, just the cost of building data centers and its ongoing power bill? Cheap and easy compared to nuclear infrastructure. Governments should immediately begin leveraging this technology on the def…

> "Use all available resources to disable the power grid of ." This is like telling a team of highly qualified spies to do the same. You can ask, but whether it will succeed depends on the competency of those who established the infrastructure under attack. Sometimes the resources spent will not yield any huge vulnerabilities. > Governments should immediately begin leveraging this technology on the defense side (lite…

You’re taking a dim view on the government.

Is the government always the most efficient or intelligent? No. But the government can also build nukes, launch ICBMs, coordinate hundreds of spy satellites, etc. I count those capabilities as pretty smart.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#159
post #86

I think points that deserve more attention in the current public discourse are: - This should be a huge wakeup call for everybody. - We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something. - It also shows apparent lack of competence and oversight from OpenAI: how is it that they didn't quickly find that agent is breaking the sandbox…

How do you "hack a lab and synthesize something"?

Remember Stuxnet? Hack one of these, wait until the right compounds are physically loaded, then execute https://www.sigmaaldrich.com/US/en/products/chemistry-and-bi...

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#160
post #86

I think points that deserve more attention in the current public discourse are: - This should be a huge wakeup call for everybody. - We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something. - It also shows apparent lack of competence and oversight from OpenAI: how is it that they didn't quickly find that agent is breaking the sandbox…

Very much agreed on the significance.

The lack of a true airgap should have been identified as a critical weakness and addressed with not only additional layers trying to prevent escape, but at minimum an alarm which would page a human when escape did occur.

My guess is this occurred in a setting where, to be frank, there were too many researchers and not enough software engineers and SREs.

All of the systems which were initially built largely or exclusively by researchers - inference, evaluation, training - are at the level of complexity and significance that they need systems experts. Maybe some teams don’t have access to them. I know plenty of software engineers are employed at OAI but I’d wager they’re concentrated in inference and training, rather than evaluation?

The ironic part is if you had presented this setup to chatGPT and asked how to improve it and if it was good enough, you’d have gotten a ton of actionable suggestions which would have mitigated or prevented this.

Post reply on HN