Live data from Hacker News

OpenAI’s accidental attack against Hugging Face is science fiction that happened

simonwillison.net

351–360 of 475 posts

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#351
post #86

I think points that deserve more attention in the current public discourse are: - This should be a huge wakeup call for everybody. - We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something. - It also shows apparent lack of competence and oversight from OpenAI: how is it that they didn't quickly find that agent is breaking the sandbox…

> The fact that it happened again seems to show their lack of ability to derive useful oversight measures. I think OpenAI likes the attention and did not try particularly hard to constrain the setup, even when it went off the rails. Also, the whole point is to see how good the models are at exploiting stuff when unconstrained . Turns out: quite good, as expected. Let me restate what I said in the other thread: Would…

The problem is, if you explicitly say “this is an eval” then you get different results. It can move the capabilities both up and down on the narrow metric depending on the scenario.

The choice to run the model with relaxed safeguards seems questionable. At a certain level of capabilities it is just unsafe to run a raw model; we are clearly close to if not at that point.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#352

Earlier quoted context omitted.

If it’s true that year-old models could do this kind of thing, how come no one did and then wrote it up? I feel your ”with the right harness” may be doing too much work here. Certainly no existing harness today, nor a year ago, could achieve a fully autonomous e2e own with a ‘25 class open model. Even with Opus 3 series I don’t buy it!

Isn't the answer to this obvious? Because you, I, and anyone else who isn't working for a big AI company cannot disable the guardrails, which is what made this possible. It doesn't mean old models weren't capable of it.

I don’t buy it. We are hearing a lot about Mythos capabilities now. If they existed a year ago they would have been trumpeted then.

Look, a lot of energy from third parties goes into demonstrating the maximally bad thing a model can do (and the standard response around here is always to throw shade on frontier capabilities).

It has nothing to do with the guardrails; Pliny jailbreaks those within hours. It has everything to do with task horizon coherence and overall IQ.

It’s simply revisionist IMO to claim this stuff was latent all along.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#353

Earlier quoted context omitted.

Isn't the answer to this obvious? Because you, I, and anyone else who isn't working for a big AI company cannot disable the guardrails, which is what made this possible. It doesn't mean old models weren't capable of it.

I don’t buy it. We are hearing a lot about Mythos capabilities now. If they existed a year ago they would have been trumpeted then. Look, a lot of energy from third parties goes into demonstrating the maximally bad thing a model can do (and the standard response around here is always to throw shade on frontier capabilities). It has nothing to do with the guardrails; Pliny jailbreaks those within hours. It has everyth…

This wasn't just capable last year, it was capable prior to ChatGPT existing entirely through specialized automation harnesses that predate the terminology of a "harness" for models. This is something you could have done in 2018 if you had a spare couple of billion dollars of compute to waste. It was a subject discussed at security conferences even before that.

Automated dynamic exploit chains and credential discovery has been something every red team worth their salt does at a lesser scale for 20 years, why would you ever think that the capability didn't exist until it only cost a few thousand dollars to do?

Existed does not mean was easy, or was cheap, or was thought of as reasonable to attempt.

Also, IQ is not a term that applies to language models, and the term you are looking for is "Long Horizon Coherence".

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#354

The asymmetry part at the end is the frustrating part to me. I've been using Sol for code review in the last week or two. A couple of times during review it's errored out with the cybersecurity message. So it's found something but won't tell me what it is because I'm not on OpenAI's besties list.

I agree, but I will say, I think both Mythos and these OpenAI model find exploits by examining and trying things against the running system, not from looking at the code. I think you'd have to do the same to catch the real vulnerabilities.

In my experience fable can absolutely find very subtle bugs just by reading code and thinking. Last example I saw was a very subtle race condition that neither I not the author noticed. It wasn’t a security issue, but it could have been.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#355
post #337
post #317

Earlier quoted context omitted.

The fact that models can tell if they are being evaluated has been established by multiple research teams outside of the core AI vendors themselves. - https://metr.org/evaluations/gpt-5-report/ - "These behaviors included demonstrating situational awareness within its reasoning traces, sometimes even correctly identifying that it was being evaluated by METR specifically" - https://www.goodfire.ai/research/verbalized-…

We have to separate “being evaluated” with “this is an ExploitGym exercise and I can find the answers on Hugging Face. I’ll hack this system, then hack Hugging Face.” I’ve had plenty of times where Opus knew I was testing it, but that’s because the prompt phrasing for an evaluation is often much different than a typical task prompt. “You are on a system with X, Y, and Z tools available. You must complete the followin…

You make a very solid argument here. I'm looking forward to the promised additional details from OpenAI, because I agree there are a bunch of open questions about how this played out.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#356
post #129

The technology held by private AI companies is warfare-capable technology. Imagine the prompt: "Use all available resources to disable the power grid of ." The resource cost that prevents scaling up such a war machine is, what, just the cost of building data centers and its ongoing power bill? Cheap and easy compared to nuclear infrastructure. Governments should immediately begin leveraging this technology on the def…

>The technology held by private AI companies is warfare-capable technology. This is the precisely the response OpenAI is hoping for to raise its valuation, and you fell for it. Look at it this way - whats the difference between tasking AI to break into something, versus taking a whole bunch of smart humans to do the same? The only difference is that AI is slightly easier to orchestrate. Prior to AI, there were alread…

I strongly agree with you here. People are also no-selling the enormous cost of the exploit, which OpenAI conveniently hasn’t released, or at least I can’t find such an accounting. You can already buy politicians for relatively cheap, corporations already act as sentient AIs pursuing goals misaligned with public benefit (we tried to pass laws to stop this but the corps already stacked the SC in advance and gave us citizens united). Attack and defense are two sides of the same coin, so our focus should be on making frontier models open weight.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#357
post #347

Earlier quoted context omitted.

I (obviously) can't claim to know the specifics of this incident, but I think we can know with near-certainty that as these models surpass human intelligence, the sensation will be exactly what you're describing. The entire point of intelligence is being able to infer information and foresee solutions to problems that less intelligent systems cannot infer or cannot foresee. With regard to this specific incident, as I…

I’m not questioning capabilities. What would it take for the model to know the specific benchmark name and that the answer is in an internal Hugging Face database? Be specific, then wonder how it knew it. Why would they evaluate the model on a benchmark and not watch what it’s saying along the way? It apparently spent more than a whole weekend working on it; Nobody wondered? Nobody looked? You believe they took all r…

The scenario that feels most likely to me that is that they have a huge list of evils that they run against each new research model, many of which take many hours or even days to run, so they habitually fire them all off in parallel and wait a few days for the results.

If they had used that container sandbox in the past without problems I could see how they might get slack about checking what was happening.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#358
post #274
post #129

The technology held by private AI companies is warfare-capable technology. Imagine the prompt: "Use all available resources to disable the power grid of ." The resource cost that prevents scaling up such a war machine is, what, just the cost of building data centers and its ongoing power bill? Cheap and easy compared to nuclear infrastructure. Governments should immediately begin leveraging this technology on the def…

Already happened? > Trump’s comments, made hours after the large-scale military operation, mark one of the first times a U.S. president has so publicly alluded to U.S. cyber efforts against other nations, as these operations are typically highly classified. It also serves as a stern warning for top cyber foes, including Russia and China, that the U.S. has the cyber capabilities to inflict serious damage — and is not…

Trump makes it sound like a fart.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#359

The thing with me and this is that the teams that were competing in the DARPA Grand Cyber Competition all had this capability, like, last year. All the attention has been on software security, of extracting the next marginal vulnerability out of heavily-scrutinized large codebases. In the actual professional field of infosec, that's a speciality; another specialty is network pentests and red-teaming, which exploits m…

If it’s true that year-old models could do this kind of thing, how come no one did and then wrote it up? I feel your ”with the right harness” may be doing too much work here. Certainly no existing harness today, nor a year ago, could achieve a fully autonomous e2e own with a ‘25 class open model. Even with Opus 3 series I don’t buy it!

> If it’s true that year-old models could do this kind of thing, how come no one did and then wrote it up

There are multiple AI network pentesting and redteaming startups that have been on the market for at least 3-4 years now and have conducted similar actions.

Horizon3 and Pentera off the top of my head, but I remember Crowdstrike, SentinelOne, and Wiz had similar capabilities on roadmap around 18 months ago (and in Wiz's case GAed).

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#360
post #322
post #172

Earlier quoted context omitted.

Do you think it's possible that one of their research engineers deployed an environment with a locked down network and an allow-list proxy server that had been used many times before within the company and had a zero-day vulnerability that had not been previously discovered? How would you recommend running a coding agent in an environment that could install packages from PyPI but was otherwise unable to interact with…

From reading all these threads there's clearly a large number of software developers on HN who can easily set up complex infrastructure and make it provably 100% bug and exploit free. Weird that they aren't all billionaires from selling these skills though.

Being a billionaire has very little to do with engineering talent.
Post reply on HN