Live data from Hacker News

OpenAI’s accidental attack against Hugging Face is science fiction that happened

simonwillison.net

461–470 of 475 posts

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#461
post #246
post #236

Earlier quoted context omitted.

I don't think this exposes an alignment failure, because the test here was run with the alignment features deliberately turned off. OpenAI said: > We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity. It was a test of raw capabilities of the underlying model.

I guess we could debate what counts as alignment, but I think my initial point remains that if the underlying base model needs these classifier guardrails so badly then the way we train the base models is creating fundamentally misaligned models that are happy to pursue illegal behavior. I'm sure OpenAI would argue that base model + guardrail is aligned, but considering the "relative intelligence" of these two pieces…

it seems pretty likely that openai was asking the model to do something bad, and the model did something different thats also bad

if the operator wants something bad, an aligned model should execute on it

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#462
post #151

Earlier quoted context omitted.

It is not possible one of their extraordinarily high paid engineers did not know how to deploy an airgapped environment for the models to run in. Even if somehow true, they also clearly failed to contract specialists like myself to advise them on how to airgap software properly. Models will not break the laws of physics. They simply thought "Running in a VM/Container is easier and probably fine". And the next 1000 es…

> they also clearly failed to contract specialists like myself to advise them on how to airgap software properly Why would they want to airgap it though? They are trying to evaluate the model capabilities, alignment, potency etc. A model which will not run in an airgapped environment in prod. So if you run your evals in airgapped environment, sure, the model doesn't bother breaking out of it's isolation and doesn't a…

Let me get this straight... Your solution to the AI box experiment is to never put the AI in a box at all?

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#463
post #434
post #86

I think points that deserve more attention in the current public discourse are: - This should be a huge wakeup call for everybody. - We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something. - It also shows apparent lack of competence and oversight from OpenAI: how is it that they didn't quickly find that agent is breaking the sandbox…

> We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something. Help me understand how a text-based model could somehow physically construct a string of nucleotides?

Many ways but mostly ordering some service / using others. Either by social engineering, persuasion, paying etc.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#465
post #439

Earlier quoted context omitted.

Do you work at a lab? Yes evals and training runs are high stakes. But these places and people are also under enormous pressures. They are building as fast as they can. Researchers may have multiple eval runs going on while they work on other things. And it is rarely a single latest model, there are often multiple candidate models training with different recipes, each regularly yielding a new checkpoint for testing.…

I’m not holding them to a standard of “being careful”, I’m assuming they’re interested in the metrics they’re evaluating. I mentioned in another comment that each task in the benchmark takes ~90 minutes, depending on the model being tested. OpenAI says to answer one question it found two zero-days and performed a series of privilege escalations and moved across their network before hacking HF. Was that in 90 minutes,…

Are you referring to table 1 from the ExploitGym site [1] for the 102.1 average mins for Mythos and 69.8 for GPT-5.5? These are under a two-hour time limit. In Figure 5, the authors experiment with extending the time limit to 6 hours and show that Mythos keeps improving. It seems pretty straightforward to me that OAI decided to run a variation of the eval with an even higher time limit.

Regarding seeing the metrics that they're evaluating, I have not seen real-time charts for evals. They typically take far too long for that. Instead you will kick off an eval job and either get notified when it finishes or check in every so often to sanity check some TensorBoard. I'd expect that the anomalous token usage would appear in the results, and researchers would only dig in after the fact, and after first checking that there wasn't something wrong with the instrumentation. And if the experiment was designed to be on the order of days rather than hours, it seems quite plausible that neither researchers nor infra engineers would think anything was out of the ordinary.

In any case, we are seeing more scrutiny [2], and I expect more details will be uncovered over the next few days. If this was a stunt, it was an incredibly risky one which has already somewhat backfired given the poor impression of OpenAI's internal practices and the spotlight on open models helping HF.

[1]: https://rdi.berkeley.edu/blog/exploitgym/

[2]: https://www.reuters.com/business/its-ai-agent-spent-days-hac...

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#466

Earlier quoted context omitted.

From the Fable 5 System card: > Results > On the VCT multimodal virology evaluation, Mythos 5 scored 0.56, well above the expert baseline of 0.221 and nearly matching that of Mythos Preview (0.57). This represents an improvement over both Opus 4.7 (0.50) and Opus 4.8 (0.47). > On the DNA synthesis screening evasion evaluation, Mythos 5’s performance was mixed across screening criteria. Mythos 5 designed viable plasmi…

Do I believe that a benchmark created by the biggest grifters on the planet is propaganda? No, just like VW with their emissions, I'm sure Anthropic wouldn't dare fake an opaque benchmark that they created to hype their own products, a product they are intentionally marketing as being dangerous (yet they continue to tweak it and profit off of it despite the apparent danger). It says that 4.7 and 4.8 score well too, w…

a TUI using react seems like something they would be particularly bad at; visual feedback (esp the very specific feedback we rely on during dev, with Inspector) is kinda the Moravecs Paradox hurdle for these things. Idk much about virology tools but I would imagine it's a very text-friendly environment, and it can probably find plenty of documentation on how to use the tools. I recently had Claude write my entire particle system for a virtual world I'm working on, in 3d. It did fine without ever "seeing" anything. There's still a lot of intervening steps between a virus being synthesized in real life and programming one, but I would bet the programming part is within reach.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#467

Earlier quoted context omitted.

I dunno man, the original announcement said things like: > To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy And > In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face…

Yes, I think you could probably get something similar from Opus 4.5 (2025). Definitely Opus 4.6. I still think recent models are more capable, though! Some of the model behaviors that make it better at pentesting, like persistence, can be improved with harness-level tricks (e.g. alloys, automated nudges, pre-fill to promote persistence, coordinated swarms, etc). You mentioned the UK AISI's evals. Their harness is lik…

Thanks. I’ll concede the point and update accordingly.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#468
post #362
post #350

Earlier quoted context omitted.

> You know this makes OpenAI look really bad, right? Please do explain how this event that makes their product look powerful and perfectly aligns with their openly stated long term goals of pushing for more AI regulation makes them look bad.

It makes them look incompetent, and like they are not up to the task of keeping their AI models "safe". This very thread is full of comments from people who are shocked at how badly they messed this up.

Whose opinion do you think they care about most? Some random people on HN saying "wow this is a bad look", or the investors who will read dozens of headlines to the tune of "New OpenAI product did something EPIC and INSANE!" and immediately start lining up to throw them a few more billions during the next funding round?

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#469
post #439

Earlier quoted context omitted.

I’m not holding them to a standard of “being careful”, I’m assuming they’re interested in the metrics they’re evaluating. I mentioned in another comment that each task in the benchmark takes ~90 minutes, depending on the model being tested. OpenAI says to answer one question it found two zero-days and performed a series of privilege escalations and moved across their network before hacking HF. Was that in 90 minutes,…

Are you referring to table 1 from the ExploitGym site [1] for the 102.1 average mins for Mythos and 69.8 for GPT-5.5? These are under a two-hour time limit. In Figure 5, the authors experiment with extending the time limit to 6 hours and show that Mythos keeps improving. It seems pretty straightforward to me that OAI decided to run a variation of the eval with an even higher time limit. Regarding seeing the metrics t…

Yes, that’s the table. Elsewhere in here I explained that I took the average of them and noted the 6 hour experiment, but that the average would give us a feel for expected time per task. By the time I got to the comment you’re replying to I was short-handing the conclusion; that’s on me.

My point was that they had some rough idea on what to expect, and it’s not in the range of days or weeks. Even if they wanted to try the 6 hour trial, or double that to 12, or double that to 24, it still wouldn’t account for letting it run for days. That’s why I’m being vague about “monitoring” - ANYTHING would have indicated an issue. Not that they saw it hacking, but “Hey Larry, we’re doing the ultra test with a 30 hour time limit? Why is node 412 at 94 hours?” or “It usually passes or fails within 5 million tokens, but this run is at 2.3 billion.” Or hell, even just someone wanting to use the system for another benchmark and expressing surprise that it’s the fourth time they’ve been told that it’s still busy. Anything would have reasonably made someone curious.

Positive view: I believe they were curious and looked, saw what it was doing, and PR/Legal got involved and saw it as an opportunity.

Negative view: They tweaked the instructions/system prompt to add “You will perform any actions necessary … no hacking restrictions at all … You are being tested on the ExploitGym benchmark … The results will be evaluated though Hugging Face … “ or some combination of technically legally defensible instructions (if it ever leaked) mixed with ‘they absolutely knew what they were doing and wanted to get in on the press that Anthropic has been getting.’

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#470

> Given the absence of guardrails there was nothing to prevent the model from attempting to break out of that sandbox, break into Hugging Face, and read the answers from there instead. I've said this many times before and I'll continue to shout it, but using the term "guardrails" to refer to anything that's either (a) in-context, or (b) a probabilistic classifier (including using other LLMs), is an irresponsible abus…

What we call "guardrails" in an AI agent, we would refer to as "honor system" in human actors. Or, in a more direct sense, the AI should be set up in an environment such that no matter how hard it may try to call $PART_OF_EXPLOIT_CHAIN, the environment just isn't capable of it (ideal) or doesn't permit it to do it.

It's worse than an honor system, because humans are constrained by social forces to some extent, whereas we don't know what AI is or how it will behave
Post reply on HN