Live data from Hacker News

OpenAI’s accidental attack against Hugging Face is science fiction that happened

simonwillison.net

311–320 of 475 posts

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#311
post #129

The technology held by private AI companies is warfare-capable technology. Imagine the prompt: "Use all available resources to disable the power grid of ." The resource cost that prevents scaling up such a war machine is, what, just the cost of building data centers and its ongoing power bill? Cheap and easy compared to nuclear infrastructure. Governments should immediately begin leveraging this technology on the def…

>The technology held by private AI companies is warfare-capable technology. This is the precisely the response OpenAI is hoping for to raise its valuation, and you fell for it. Look at it this way - whats the difference between tasking AI to break into something, versus taking a whole bunch of smart humans to do the same? The only difference is that AI is slightly easier to orchestrate. Prior to AI, there were alread…

I find myself in major disagreement here. The nice thing about humans is we always have context and ongoing internal conversations including about ethics. If you recruit a bunch of hackers to take down a country, not only is the pay an order of magnitude higher, you have to worry about them backstabbing you, leaking your intent to the government, whistleblowing to the press, and so on. It’s not trivial to do that with a group of (especially capable) humans. They will also have differences of opinion with you and coworkers with some regularity.

I try to recruit a bunch of people to attack a country and it’s going to be hard to get people to say yes, and they will definitely ask or find out which country, and wonder about potential retribution. You see this dynamic show up even to some extent among cybergangs, not all targets are equal.

A single private individual wielding a compliant and hyper capable LLM is an entirely different paradigm. They are accountable to nearly no one and often have few brakes. Frequently they may not care about avoiding detection. And the AI itself may be incapable of the same scale of self reflection and brake behavior a human team will.

We may potentially be entering the age of lone wolf cyberterrorism, and some of the same principles and problems apply. When it is easier for single people to plot and carry out high-impact, destructive acts they happen more often. Doubly so if there’s a social contagion. Gun violence isn’t actually a bad analogy here. And do you remember how many corporate sites got defaced in the prime Anon era?

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#312
post #252

Earlier quoted context omitted.

I want a model that can find every security vulnerability in the software I write, including crafting POC exploits against those vulnerabilities so I can be absolutely sure that I have fixed them. A model that can do that is aligned with me. The unsolveable problem is a model that can tell the difference between me saying "I wrote this software and need you to find vulnerabilities" when it's TRUE v.s. me saying the e…

I find this fascinating because - let's say that I want to simulate the scenario for myself that a language model finds itself in. I'll put aside my deep dislike of anthropomorphizing AI for a minute. So if I'm an LLM, the only "sense" that I have available to me is the incoming stream of tokens. I can emulate that by forcing myself to imagine evaluating incoming requests by putting myself in a completely empty room…

> In the "real world" of course, I have five senses to rely upon. Critically, I have the ability to collect additional context "out of band" of the conversation and interact with unrelated entities to either confirm or refute claims.

The problem with thinking about LLM context is that we compare it to our own sense of context, which is much larger and continually expands and revises itself while online.

I can rent human-level context for twenty bucks an hour, and the agent instructions I'm already playing around with are also the onboarding docs I should have created years ago. Truly useful agent orchestration/routing would come with Craigslist or Fiver integration.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#313
post #276

Earlier quoted context omitted.

> they also clearly failed to contract specialists like myself to advise them on how to airgap software properly Why would they want to airgap it though? They are trying to evaluate the model capabilities, alignment, potency etc. A model which will not run in an airgapped environment in prod. So if you run your evals in airgapped environment, sure, the model doesn't bother breaking out of it's isolation and doesn't a…

You can simulate the internet in an airgapped environment for the tests and services you want it to interact with if you have enough disk space, and, they absolutely do.

Sure, I'm not saying it's an untractable problem. But it's not as dumb as "how come a team of engineers paid millions haven't even heard of airgapping".

Creating a fake internet-like environment good enough to trick an advanced model that is very good at finding intricate flaws is not a simple job. Especially given that models can decide to behave differently once they suspect they might be under evaluation [1], which a fake internet would absolutely give away.

[1]: https://www.anthropic.com/research/alignment-faking

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#314
post #276

Earlier quoted context omitted.

You can simulate the internet in an airgapped environment for the tests and services you want it to interact with if you have enough disk space, and, they absolutely do.

Sure, I'm not saying it's an untractable problem. But it's not as dumb as "how come a team of engineers paid millions haven't even heard of airgapping". Creating a fake internet-like environment good enough to trick an advanced model that is very good at finding intricate flaws is not a simple job. Especially given that models can decide to behave differently once they suspect they might be under evaluation [1], whic…

I am not saying responsible research is easy, but I am saying that at their budget and scale there are no excuses to cut major corners on safety when the stakes and potential for very bad outcomes for others are this high.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#315
post #302

Agreed that people claiming “marketing stunt” need to pull their heads out the sand, but likewise Simon needs to do some of his own beach-cranium-dislodging for laying the blame of constraints on the US govt. Before the export controls were ever floated, Glasswind found many thousands of exploits, and offered patches/fixes for approximately none of them. (perhaps their exploit capability far outstrips their remediati…

Huh, I thought Anthropic had been offering patches as well as reports, but their tracker at https://red.anthropic.com/2026/cvd/ledger/ lists 1,596 disclosures and currently shows only 27 of those as fixed.

But https://www.anthropic.com/coordinated-vulnerability-disclosu... says:

> Every report we send generally reflects a finding that a human security researcher has reviewed and confirmed. Reports originating from AI-powered discovery are clearly labeled as such. Where we have access to source and our tooling produces a potential candidate patch, we include it, labeled by provenance and offer to collaborate with the maintainer on a production-quality fix.

So I'm not sure why so few of the reported issues have a confirmed patch.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#316
post #118
post #115

Earlier quoted context omitted.

I think there's a reasonable case that the agent knew it was breaking into the system, for some definition of knew . I think OpenAI would be very reluctant to let this go to a place where the reasoning was part of discovery.

Agents aren't subjects of criminal law. I agree there may be civil liability, I know far less about that.

Ok, but if the agent's reasoning log says "The best way to get into Hugging Face is to find and exploit a zero-day vulnerability", surely those responsible for monitoring its actions should be criminally liable.

These guys would be screwed if they were operating under the EU AI Act.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#317
post #309

Earlier quoted context omitted.

1. It's been extremely well established for multiple generations of models that they have no problem detecting when they're being evaluated. Pretty sure it was Opus 4.8 that the independent evaluators literally filed an assessment that said "We have no assessment to make as [model] consistently detected it was being evaluated, making our assessments untrustworthy." 2. Regardless of whether the model was being watched…

1. It has not been established, it has been stated by the company that has a strong motive to make their “intelligence product” sound almost otherworldly. That motivation is the basis for my suspicion. 2. Hah no… that’d be silly. I mean watching it like you might watch Claude Code or literally any other AI interface. Literally just be in the area watching what it outputs. Again, they’re text based. You don’t have to…

The fact that models can tell if they are being evaluated has been established by multiple research teams outside of the core AI vendors themselves.

- https://metr.org/evaluations/gpt-5-report/ - "These behaviors included demonstrating situational awareness within its reasoning traces, sometimes even correctly identifying that it was being evaluated by METR specifically"

- https://www.goodfire.ai/research/verbalized-eval-awareness-i... - "We study verbalized eval awareness — cases where a model organically expresses awareness of being evaluated — and show that it appears across more models and benchmarks than previously documented."

- https://www.apolloresearch.ai/science/more-capable-models-ar... - "We expect that the most capable models increasingly realize that they are being evaluated, and this reduces the utility of the evals."

- https://www.aisi.gov.uk/blog/evaluating-whether-ai-models-wo... - "An important limitation of our work is evaluation awareness, where models recognise that they are being evaluated, which may lead them to alter their behaviour and thereby undermine the reliability of our results. We found that all models we tested can reliably distinguish our evaluation scenarios from deployment data when prompted."

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#318
post #172
post #151

Earlier quoted context omitted.

It is not possible one of their extraordinarily high paid engineers did not know how to deploy an airgapped environment for the models to run in. Even if somehow true, they also clearly failed to contract specialists like myself to advise them on how to airgap software properly. Models will not break the laws of physics. They simply thought "Running in a VM/Container is easier and probably fine". And the next 1000 es…

Do you think it's possible that one of their research engineers deployed an environment with a locked down network and an allow-list proxy server that had been used many times before within the company and had a zero-day vulnerability that had not been previously discovered? How would you recommend running a coding agent in an environment that could install packages from PyPI but was otherwise unable to interact with…

> How would you recommend running a coding agent in an environment that could install packages from PyPI but was otherwise unable to interact with the wider world?

You download the "website" and make it available locally. That's trivial.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#319

Earlier quoted context omitted.

How does it prove that?

Do you think sama deliberately attacked HuggingFace and then claimed it was a rogue model? OpenAI is one of the most scrutinized companies in the world right now. Sam's house was independently firebombed and then shot at 3 months ago. HuggingFace is a foreign competitor with every incentive to call out foul play from American frontier labs. Why flagrantly break the law and invite investigation just for a PR moment wh…

I'm not saying that the conspiracy theory is true, but I do want to note that Greg Brockman, co-founder and President of OpenAI, is an angel investor in HuggingFace. As companies, they are not totally unconnected and opposed.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#320

Earlier quoted context omitted.

Because this is amazing PR? Just following the Anthropic rulebook.

Same rulebook whereby they shot themselves in the foot and got their model clipped by the government?

And then everyone wanted to pay for the model that was so smart it was banned?
Post reply on HN