Live data from Hacker News

OpenAI’s accidental attack against Hugging Face is science fiction that happened

simonwillison.net

271–280 of 475 posts

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#271
post #256
post #252

Earlier quoted context omitted.

I want a model that can find every security vulnerability in the software I write, including crafting POC exploits against those vulnerabilities so I can be absolutely sure that I have fixed them. A model that can do that is aligned with me. The unsolveable problem is a model that can tell the difference between me saying "I wrote this software and need you to find vulnerabilities" when it's TRUE v.s. me saying the e…

But in the process of finding every security vulnerability in the software you write, would you be ok with your model hacking AWS to start mining bitcoin? Would that still be aligned with you? (I'm guessing not) That's the alignment problem I'm referring to (which is one of the many aspects of alignment), for which we do not have robust recipes, and not only that but for which research suggests it is becoming harder…

I don't know, but these were exactly the questions the industry had to handle with CORE Impact and Immunity Canvas, and ultimately all the way back to Dan Farmer's SATAN before that.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#272

Earlier quoted context omitted.

> Russians and Chinese [...] are using their propaganda machines to make Americans hate datacenters Are they though? or is that the story the people most invested in ai have an interest in making us believe? [0] ... `it's the foreign bad guys propaganda machines making you believe your city struggling for water and electricity is a bad thing`. People as a whole may not always be the brightest, but threaten their imme…

See u/ verdverm's response, below

yes and see: https://text.npr.org/nx-s1-5844328 and Occam's razor.

Russia and China are gonna Russia and China, every issue is going to have a measure of outside influence, thats just how it is now.

But them directing large scale operation forces focused on something that is genuinely bad for the people in the area of datacenters (ie, an issue that will take care of itself from within) instead of using their resources on other issues that actually need a real propaganda push? Versus the benefit of those invested in building the datacenters using fear to get people to focus away from the resources being diverted from them to datacenters?

One side has reality working for what they want, while another side needs the propaganda to turn their billions into trillions.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#273
post #194

Earlier quoted context omitted.

It proves that it was certainly not a sama PR stunt, as many cynics are alleging.

How does it prove that?

Do you think sama deliberately attacked HuggingFace and then claimed it was a rogue model?

OpenAI is one of the most scrutinized companies in the world right now. Sam's house was independently firebombed and then shot at 3 months ago. HuggingFace is a foreign competitor with every incentive to call out foul play from American frontier labs. Why flagrantly break the law and invite investigation just for a PR moment which is already backfiring in favor of open models?

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#274
post #129

The technology held by private AI companies is warfare-capable technology. Imagine the prompt: "Use all available resources to disable the power grid of ." The resource cost that prevents scaling up such a war machine is, what, just the cost of building data centers and its ongoing power bill? Cheap and easy compared to nuclear infrastructure. Governments should immediately begin leveraging this technology on the def…

Already happened?

> Trump’s comments, made hours after the large-scale military operation, mark one of the first times a U.S. president has so publicly alluded to U.S. cyber efforts against other nations, as these operations are typically highly classified. It also serves as a stern warning for top cyber foes, including Russia and China, that the U.S. has the cyber capabilities to inflict serious damage — and is not shy about using them.

> “Policymakers are getting more comfortable employing and, crucially, acknowledging cyber operations as tools of statecraft and military power,” said Michael Sulmeyer, former assistant secretary of Defense for cyber policy under the Biden administration. “It is one thing to do it; it is another to say it.”

> The Jan. 3 strikes on Venezuela’s capital and subsequent seizure of Maduro and his wife involved close coordination among federal agencies and military units, and took months of careful planning. In a press conference following the strikes, Caine said U.S. Cyber Command, U.S. Space Command and other combatant commands “began layering different effects” to “create a pathway” for U.S. forces flying into the country before dawn Saturday.

> Trump, at the same press conference, was more overt in his description of U.S. cyber involvement: “The lights of Caracas were largely turned off due to a certain expertise that we have,” he said. “It was dark, and it was deadly.”

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#275
post #259

If by "accidental" you mean "humans deliberately trained the AI to do that", then ok. IMO, this was a PR stunt to goad the Feds into regulating AI to shore up OpenAI's moat against open source models.

Given Anthropic lost two weeks of peak Fable 5 sales to a US government restriction (and by the time they could sell it again OpenAI's GPT-5.6 had taken some wind out of its sails) I would hope that the big AI labs have learned that goading the Feds can backfire spectacularly.

You’re viewing it from a freedom not a profit perspective. If a license to use AI is required, that creates artificial scarcity. The price goes up. If gov declares Qwen et.al apostate, lack of competition increases prices.

And Anthropic’s delayed rollout was a direct response to them trying to impose extra-legislative rules on The Pentagon. I kinda doubt OpenAI has such ‘scruples’.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#276
post #151

Earlier quoted context omitted.

It is not possible one of their extraordinarily high paid engineers did not know how to deploy an airgapped environment for the models to run in. Even if somehow true, they also clearly failed to contract specialists like myself to advise them on how to airgap software properly. Models will not break the laws of physics. They simply thought "Running in a VM/Container is easier and probably fine". And the next 1000 es…

> they also clearly failed to contract specialists like myself to advise them on how to airgap software properly Why would they want to airgap it though? They are trying to evaluate the model capabilities, alignment, potency etc. A model which will not run in an airgapped environment in prod. So if you run your evals in airgapped environment, sure, the model doesn't bother breaking out of it's isolation and doesn't a…

You can simulate the internet in an airgapped environment for the tests and services you want it to interact with if you have enough disk space, and, they absolutely do.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#277
post #144

Currently trying to avoid an open weight model ban while OpenAI, who is closed, lets theirs run wild on the internet causing harm to another company because they do not understand how to airgap things. Cool.

That's the fun part. They do know how to airgap things. The models are outsmarting already pretty smart people!

I am the author of AirgapOS and I have designed systems to run in underground zero emissions chambers that are interacted with via carefully verified sd cards and/or fiber optic serial terminals.

If OpenAI had done this, the attack would not have happened. In high risk computation 0days must be in your threat model from the start, so you secure things with the laws of physics.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#278

I'm not skeptical that this attack happened, I'm skeptical that the model's prompt was truly just "solve this benchmark" and nothing more. I'm also trying to figure out why OpenAI put out a press release about this. In what way is this not admitting to a federal crime?

Because this is amazing PR? Just following the Anthropic rulebook.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#279
post #259

Earlier quoted context omitted.

Given Anthropic lost two weeks of peak Fable 5 sales to a US government restriction (and by the time they could sell it again OpenAI's GPT-5.6 had taken some wind out of its sails) I would hope that the big AI labs have learned that goading the Feds can backfire spectacularly.

You’re viewing it from a freedom not a profit perspective. If a license to use AI is required, that creates artificial scarcity. The price goes up. If gov declares Qwen et.al apostate, lack of competition increases prices. And Anthropic’s delayed rollout was a direct response to them trying to impose extra-legislative rules on The Pentagon. I kinda doubt OpenAI has such ‘scruples’.

I think OpenAI are smart enough to have looked at the Fable situation and decided that, given the unpredictable nature of the current administration, stunts like deliberately hacking another company and pretending that it was an autonomous agents gone wrong are not worth the risk.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#280
post #252
post #246

Earlier quoted context omitted.

I guess we could debate what counts as alignment, but I think my initial point remains that if the underlying base model needs these classifier guardrails so badly then the way we train the base models is creating fundamentally misaligned models that are happy to pursue illegal behavior. I'm sure OpenAI would argue that base model + guardrail is aligned, but considering the "relative intelligence" of these two pieces…

I want a model that can find every security vulnerability in the software I write, including crafting POC exploits against those vulnerabilities so I can be absolutely sure that I have fixed them. A model that can do that is aligned with me. The unsolveable problem is a model that can tell the difference between me saying "I wrote this software and need you to find vulnerabilities" when it's TRUE v.s. me saying the e…

As humans we are also vulnerable to that class of attack. A blue amazon vest and some boxes would get you into most buildings
Post reply on HN