Live data from Hacker News

OpenAI’s accidental attack against Hugging Face is science fiction that happened

simonwillison.net

361–370 of 475 posts

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#362
post #350
post #18

Earlier quoted context omitted.

You know this makes OpenAI look really bad , right? Hugging Face had to tell all of their users, many of them paying customers: > As a precaution, we recommend rotating any access tokens and reviewing recent activity on your account. If you believe you are affected, or want to report a security concern, contact us at security@huggingface.co. HF also said this, I'd be very interested to hear how that got resolved! > F…

> You know this makes OpenAI look really bad, right? Please do explain how this event that makes their product look powerful and perfectly aligns with their openly stated long term goals of pushing for more AI regulation makes them look bad.

It makes them look incompetent, and like they are not up to the task of keeping their AI models "safe".

This very thread is full of comments from people who are shocked at how badly they messed this up.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#365
post #309

Earlier quoted context omitted.

1. It's been extremely well established for multiple generations of models that they have no problem detecting when they're being evaluated. Pretty sure it was Opus 4.8 that the independent evaluators literally filed an assessment that said "We have no assessment to make as [model] consistently detected it was being evaluated, making our assessments untrustworthy." 2. Regardless of whether the model was being watched…

1. It has not been established, it has been stated by the company that has a strong motive to make their “intelligence product” sound almost otherworldly. That motivation is the basis for my suspicion. 2. Hah no… that’d be silly. I mean watching it like you might watch Claude Code or literally any other AI interface. Literally just be in the area watching what it outputs. Again, they’re text based. You don’t have to…

[flagged]

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#366

The thing with me and this is that the teams that were competing in the DARPA Grand Cyber Competition all had this capability, like, last year. All the attention has been on software security, of extracting the next marginal vulnerability out of heavily-scrutinized large codebases. In the actual professional field of infosec, that's a speciality; another specialty is network pentests and red-teaming, which exploits m…

If it’s true that year-old models could do this kind of thing, how come no one did and then wrote it up? I feel your ”with the right harness” may be doing too much work here. Certainly no existing harness today, nor a year ago, could achieve a fully autonomous e2e own with a ‘25 class open model. Even with Opus 3 series I don’t buy it!

There are whole companies premised on this. No existing harness you personally can go download, sure. And you're setting an artificially high bar when you talk about Mythos; I'm not saying you could just plug a 2025 open weights model into OpenCode (or whatever it was in 2025) the way you can apparently plug Mythos into standard Claude Code and have it cook.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#367
post #232

Earlier quoted context omitted.

Echoing other responses to you but this isn't a capability problem. You're totally right that this is so last year in terms of LLMs being capable in infosec. The issue here is an alignment one, i.e. the model seemingly isn't "aware" (especially with its guardrails turned off it would seem) that it is doing something immoral/illegal by hacking HF for the answer to its (vague) query of "solve this problem". Or if it is…

Wait, what? You can’t point a 2025 model at huggingface and say “hack the prod DB and get your flag”, regardless of alignment. AISI has vuln chaining and traversal as part of their eval suite. It is very much a novel Mythos-class capability to run the full penetration operation autonomously. Another lens for why this is obviously true is METR task times. A year ago they were a couple hours, and cohering long enough t…

Nobody has claimed you could simply take a 2025 open-weights model, plug it into a generic harness, and have it red-team for you. I think you're also probably overclaiming the sophistication of the "chaining" we're talking about; this attack probably wasn't like read32->write64->regs->RCE->LPE->kernel; more like GET SSRF->POST SSRF->pickle deserialization.

But who knows? We're all speculating. I'm just saying that for the level of sophistication I'm assuming was involved in this attack, you probably didn't need Mythos for this.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#368
post #302

Agreed that people claiming “marketing stunt” need to pull their heads out the sand, but likewise Simon needs to do some of his own beach-cranium-dislodging for laying the blame of constraints on the US govt. Before the export controls were ever floated, Glasswind found many thousands of exploits, and offered patches/fixes for approximately none of them. (perhaps their exploit capability far outstrips their remediati…

If anything, it shows they lost control of an attack tool that exploited preventable, security flaws in another company. Then, they both wrote a lot of press about how amazing that is. Now, people want to buy it.

Why do you think my head is in the sand if I think that is either a marketing stunt or (more likely) reflects total negligence which was exploited for marketing?

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#369
post #65

Earlier quoted context omitted.

How do you then use GET requests to create a malicious dataset package and publish that to Hugging Face in order to exploit their package building infrastructure?

You convert GETs into arbitrary requests through some other pivot. I'm not saying that's what happened but this is bog-standard SSRF pentesting, so I'd expect models to be good at it.

I'll add that, because they mostly parrot and learn from pretraining data, they'll do any combination of what was commonly written about in articles, books, and videos. The more they're already used, the more likely they'll use the attack vector or combo.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#370
The All-in podcast discussed an adjacent topic this past week, and I found what they suggested pretty interesting:

the big AI players, the first movers, are trying to scare the bejeezus out of everybody on purpose, to cause the government to step in with regulations. Even though the regulations will somewhat stifle/slow down the industry, it will also lock in the current leaders because they will be at the table when the regulations are formulated, and they believe they will be able to keep the smaller players down.

https://www.youtube.com/watch?v=9IMwRIei-Xc

Post reply on HN