Live data from Hacker News

OpenAI’s accidental attack against Hugging Face is science fiction that happened

simonwillison.net

251–260 of 475 posts

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#252
post #246
post #236

Earlier quoted context omitted.

I don't think this exposes an alignment failure, because the test here was run with the alignment features deliberately turned off. OpenAI said: > We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity. It was a test of raw capabilities of the underlying model.

I guess we could debate what counts as alignment, but I think my initial point remains that if the underlying base model needs these classifier guardrails so badly then the way we train the base models is creating fundamentally misaligned models that are happy to pursue illegal behavior. I'm sure OpenAI would argue that base model + guardrail is aligned, but considering the "relative intelligence" of these two pieces…

I want a model that can find every security vulnerability in the software I write, including crafting POC exploits against those vulnerabilities so I can be absolutely sure that I have fixed them.

A model that can do that is aligned with me.

The unsolveable problem is a model that can tell the difference between me saying "I wrote this software and need you to find vulnerabilities" when it's TRUE v.s. me saying the exact same thing and setting it loose on software written by other people where my intent is to exploit that software (and not to report the issues to them.)

Even AGI doesn't give you a model that can read minds and forecast the future.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#253
post #236

Earlier quoted context omitted.

I don't think this exposes an alignment failure, because the test here was run with the alignment features deliberately turned off. OpenAI said: > We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity. It was a test of raw capabilities of the underlying model.

Alignment isn't alignment if it can be turned on and off at the whim of company employees. This time the damage was minor, relatively speaking. What happens when a model just "testing its capabilities" breaks into banking infrastructure or government military assets? The damage could be catastrophic.

Alignment with who in what context? Likely an unresolvable debate like consciousness, where there is not single or right answer for everyone.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#254

This makes me think: what are the odds weights from Frontier Models have already been stolen? Something like Mythos without "guardrails" seems like a hell of a glittering gem for many nefarious actors.

i mentioned this in another comment but if these models are fully capable by default and not neutered in the weights themselves via training then you can assume they've already been stolen. It would be worth any cost to steal it.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#255

Earlier quoted context omitted.

I believe the Russians and Chinese recognized this years ago, which is why they are using their propaganda machines to make Americans hate datacenters.

Is there a source? Or is that xenophobia?

https://www.nytimes.com/2026/07/09/business/china-russia-ai-...

https://archive.ph/yAvgz

There is no doubt an amount of xeno/sinophobia is at play as well. Russia is still killing innocent people in an illegal invasion of Ukraine, they earned their keep

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#256
post #252
post #246

Earlier quoted context omitted.

I guess we could debate what counts as alignment, but I think my initial point remains that if the underlying base model needs these classifier guardrails so badly then the way we train the base models is creating fundamentally misaligned models that are happy to pursue illegal behavior. I'm sure OpenAI would argue that base model + guardrail is aligned, but considering the "relative intelligence" of these two pieces…

I want a model that can find every security vulnerability in the software I write, including crafting POC exploits against those vulnerabilities so I can be absolutely sure that I have fixed them. A model that can do that is aligned with me. The unsolveable problem is a model that can tell the difference between me saying "I wrote this software and need you to find vulnerabilities" when it's TRUE v.s. me saying the e…

But in the process of finding every security vulnerability in the software you write, would you be ok with your model hacking AWS to start mining bitcoin? Would that still be aligned with you? (I'm guessing not)

That's the alignment problem I'm referring to (which is one of the many aspects of alignment), for which we do not have robust recipes, and not only that but for which research suggests it is becoming harder to create guardrails for as base models get smarter.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#258
post #129

The technology held by private AI companies is warfare-capable technology. Imagine the prompt: "Use all available resources to disable the power grid of ." The resource cost that prevents scaling up such a war machine is, what, just the cost of building data centers and its ongoing power bill? Cheap and easy compared to nuclear infrastructure. Governments should immediately begin leveraging this technology on the def…

I believe the Russians and Chinese recognized this years ago, which is why they are using their propaganda machines to make Americans hate datacenters.

> Russians and Chinese [...] are using their propaganda machines to make Americans hate datacenters

Are they though? or is that the story the people most invested in ai have an interest in making us believe? [0]

... `it's the foreign bad guys propaganda machines making you believe your city struggling for water and electricity is a bad thing`. People as a whole may not always be the brightest, but threaten their immediate survival needs (ie; not some vague climate change most people don't understand or see, but) actual power outages, water and rolling blackouts -- people will quickly pay attention and care.

[0] https://text.npr.org/nx-s1-5844328

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#259

If by "accidental" you mean "humans deliberately trained the AI to do that", then ok. IMO, this was a PR stunt to goad the Feds into regulating AI to shore up OpenAI's moat against open source models.

Given Anthropic lost two weeks of peak Fable 5 sales to a US government restriction (and by the time they could sell it again OpenAI's GPT-5.6 had taken some wind out of its sails) I would hope that the big AI labs have learned that goading the Feds can backfire spectacularly.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#260

Earlier quoted context omitted.

Are you serious? Before LLM’s you needed serious skills and experience to pull this off. Now it’s a prompt away on some terminal done by any random dud. And I dont mention the velocity of iteration or that they will be even better in 1 year.

Its like you don't even read the post or the content. Hint: agentic loops. No its not a prompt away.

Also the harness, there have been several HN submissions showing that older/smaller models can find a subset of the exploits given a good env
Post reply on HN