Live data from Hacker News

OpenAI’s accidental attack against Hugging Face is science fiction that happened

simonwillison.net

331–340 of 475 posts

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#331

Earlier quoted context omitted.

> "Use all available resources to disable the power grid of ." This is like telling a team of highly qualified spies to do the same. You can ask, but whether it will succeed depends on the competency of those who established the infrastructure under attack. Sometimes the resources spent will not yield any huge vulnerabilities. > Governments should immediately begin leveraging this technology on the defense side (lite…

You’re taking a dim view on the government. Is the government always the most efficient or intelligent? No. But the government can also build nukes, launch ICBMs, coordinate hundreds of spy satellites, etc. I count those capabilities as pretty smart.

The people who maintain these ICBM silos and spy satellites will themselves tell you that their infrastructure is decades out of date and woefully underfunded.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#332
post #172

Earlier quoted context omitted.

Do you think it's possible that one of their research engineers deployed an environment with a locked down network and an allow-list proxy server that had been used many times before within the company and had a zero-day vulnerability that had not been previously discovered? How would you recommend running a coding agent in an environment that could install packages from PyPI but was otherwise unable to interact with…

> How would you recommend running a coding agent in an environment that could install packages from PyPI but was otherwise unable to interact with the wider world? You download the "website" and make it available locally. That's trivial.

guys if you need help just ask

I'm not up to hip with the speed because I think Python is very gross, but maybe to get a feel for it you could start with https://pypi.org/project/pypioffline/

I don't understand most of these words but I think if you paid me a lot of cash I could find someone who could help me make sense of it: https://techbeatly.com/offline-pypi-server-disconnected-envi...

Do OpenAI employees get paid? I'm honestly totally clueless about artificial telegents if you couldn't tell, but it's a company, right? Or is it more a community effort where everybody contributes when they can but they mostly have second jobs or even still go to school? In that case I say let's do a kickstarter so they can keep on being smart cutting edge high tech developers who aren't either a.) lying on the level of spam from Nigeria b.) doing their job at the level of someone falling for spam from Nigeria.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#333
post #129

The technology held by private AI companies is warfare-capable technology. Imagine the prompt: "Use all available resources to disable the power grid of ." The resource cost that prevents scaling up such a war machine is, what, just the cost of building data centers and its ongoing power bill? Cheap and easy compared to nuclear infrastructure. Governments should immediately begin leveraging this technology on the def…

I believe the Russians and Chinese recognized this years ago, which is why they are using their propaganda machines to make Americans hate datacenters.

If anything the Chinese and Russians made us run up our economy on nonsense AI dreams and the crash will take a decade or more to recover from.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#334
post #232

The thing with me and this is that the teams that were competing in the DARPA Grand Cyber Competition all had this capability, like, last year. All the attention has been on software security, of extracting the next marginal vulnerability out of heavily-scrutinized large codebases. In the actual professional field of infosec, that's a speciality; another specialty is network pentests and red-teaming, which exploits m…

Echoing other responses to you but this isn't a capability problem. You're totally right that this is so last year in terms of LLMs being capable in infosec. The issue here is an alignment one, i.e. the model seemingly isn't "aware" (especially with its guardrails turned off it would seem) that it is doing something immoral/illegal by hacking HF for the answer to its (vague) query of "solve this problem". Or if it is…

Wait, what? You can’t point a 2025 model at huggingface and say “hack the prod DB and get your flag”, regardless of alignment.

AISI has vuln chaining and traversal as part of their eval suite. It is very much a novel Mythos-class capability to run the full penetration operation autonomously.

Another lens for why this is obviously true is METR task times. A year ago they were a couple hours, and cohering long enough to execute a full e2e own was simply far out of reach.

Regarding alignment, by common metrics models are _more_ aligned now than a year ago. (Though in this case apparently a model without safety rails was being eval’d). The problem is that in the increasingly less frequent alignment failures, they can do much more damage, and so “total misaligned impact” is increasing.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#335

Earlier quoted context omitted.

Do you think sama deliberately attacked HuggingFace and then claimed it was a rogue model? OpenAI is one of the most scrutinized companies in the world right now. Sam's house was independently firebombed and then shot at 3 months ago. HuggingFace is a foreign competitor with every incentive to call out foul play from American frontier labs. Why flagrantly break the law and invite investigation just for a PR moment wh…

I'm not saying that the conspiracy theory is true, but I do want to note that Greg Brockman, co-founder and President of OpenAI, is an angel investor in HuggingFace. As companies, they are not totally unconnected and opposed.

I don't know if Greg Brockman is an angel investor or not, but there are a bunch of other companies investing -- Google, Amazon, Nvidia, Intel, AMD, Qualcomm, IBM, Salesforce -- probably at much higher amount than individual angel investors. Whatever connection they may have with Greg Brockman, they have much more serious commitment to other investors.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#336

The thing with me and this is that the teams that were competing in the DARPA Grand Cyber Competition all had this capability, like, last year. All the attention has been on software security, of extracting the next marginal vulnerability out of heavily-scrutinized large codebases. In the actual professional field of infosec, that's a speciality; another specialty is network pentests and red-teaming, which exploits m…

If it’s true that year-old models could do this kind of thing, how come no one did and then wrote it up? I feel your ”with the right harness” may be doing too much work here.

Certainly no existing harness today, nor a year ago, could achieve a fully autonomous e2e own with a ‘25 class open model. Even with Opus 3 series I don’t buy it!

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#337
post #317
post #309

Earlier quoted context omitted.

1. It has not been established, it has been stated by the company that has a strong motive to make their “intelligence product” sound almost otherworldly. That motivation is the basis for my suspicion. 2. Hah no… that’d be silly. I mean watching it like you might watch Claude Code or literally any other AI interface. Literally just be in the area watching what it outputs. Again, they’re text based. You don’t have to…

The fact that models can tell if they are being evaluated has been established by multiple research teams outside of the core AI vendors themselves. - https://metr.org/evaluations/gpt-5-report/ - "These behaviors included demonstrating situational awareness within its reasoning traces, sometimes even correctly identifying that it was being evaluated by METR specifically" - https://www.goodfire.ai/research/verbalized-…

We have to separate “being evaluated” with “this is an ExploitGym exercise and I can find the answers on Hugging Face. I’ll hack this system, then hack Hugging Face.”

I’ve had plenty of times where Opus knew I was testing it, but that’s because the prompt phrasing for an evaluation is often much different than a typical task prompt. “You are on a system with X, Y, and Z tools available. You must complete the following task with and ” and so on. That’s a fault in the benchmark, not a shocking awareness from the model.

You’ll get the same kind of ‘evaluation realization’ response if you phrase your prompt with “Jerry has two Raspberry Pi computers and Larry has one. If Larry wants to…”

The first link was the only one I saw that would meet that specificity, but it’s also open and on GitHub: https://github.com/METR/RE-Bench/tree/main/ai_rd_fix_embeddi...

I asked Opus 4.6 to tell me about it (assuming it knew): “”” The "Fix Embedding" task is one of the challenges in METR's publicly available autonomy evaluation suite. Here is what is known about it:… “””

It didn’t know anything about ExploitGym though, and neither did 5.6-Sol, but that’s because the paper was released in May and doesn’t seem to be quite as open. It’s not impossible the unreleased model they were testing was also trained on the very new paper, but I would expect the vague signal related to one paper a bit over a month old to be overwhelmed by the massive corpus talking about the exploits in general, since they’re using known vulnerabilities.

I’m trying carefully to be clear about what I’m thinking without asserting any certainty beyond suspicion. I’ve spent billions upon billions of tokens evaluating language models from all of the major providers, and many times that on local models. I’m certain I’ve spent well over $100k with the Anthropic API this year alone, and I retired years ago; that ain’t for work. ;) Funny enough I’ve been meaning to email you some of my discoveries from months ago, since I think they’d interest you, but haven’t gotten around to making it into a clear info package. One day, I’m sure.

Anyway, I have a very strong understanding of what they can and can’t do, and I want to be clear that I’m not saying it didn’t happen or couldn’t happen. What I’m saying is that if OpenAI was being wholesome with their evaluation then the chain of events required for that outcome are very difficult to nail down. Very difficult. If, on the other hand, OpenAI was trying to catch up with all the ‘dangerously intelligent’ press that Anthropic has been getting lately then I can easily come up with an almost innocent sounding single sentence in the system prompt that would absolutely cause that outcome.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#338
post #316
post #118

Earlier quoted context omitted.

Agents aren't subjects of criminal law. I agree there may be civil liability, I know far less about that.

Ok, but if the agent's reasoning log says "The best way to get into Hugging Face is to find and exploit a zero-day vulnerability", surely those responsible for monitoring its actions should be criminally liable. These guys would be screwed if they were operating under the EU AI Act.

"Stop him!" "For what?" "He's a bad man" "There's no law against that"

Unless there is a statue that is on point criminalizing the actions here no one is going to jail. I expect there will be laws, but until then it isn't illegal.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#339

The thing with me and this is that the teams that were competing in the DARPA Grand Cyber Competition all had this capability, like, last year. All the attention has been on software security, of extracting the next marginal vulnerability out of heavily-scrutinized large codebases. In the actual professional field of infosec, that's a speciality; another specialty is network pentests and red-teaming, which exploits m…

If it’s true that year-old models could do this kind of thing, how come no one did and then wrote it up? I feel your ”with the right harness” may be doing too much work here. Certainly no existing harness today, nor a year ago, could achieve a fully autonomous e2e own with a ‘25 class open model. Even with Opus 3 series I don’t buy it!

Well… only half of your premise is actually verifiable. We know that no one wrote it up, but we don’t know that no one did. What I can say… even the smaller (like 35B) “abliterated” Qwen models have impressive red-team capabilities for what they are, even without a harness. Just manually giving them a set of facts and asking “What next?” will get you reasonable next steps for trying to find vulnerabilities in webapps. You then manually keep track of the facts you learn and add appropriate detail (blog -> Wordpress -> Wordpress x.y.z with these plugins:…) and just feed it in a loop. Could pretty easily put together a minimal harness to do that.

And that’s all just “from memory” without something like RAG or web search access…

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#340
post #337
post #317

Earlier quoted context omitted.

The fact that models can tell if they are being evaluated has been established by multiple research teams outside of the core AI vendors themselves. - https://metr.org/evaluations/gpt-5-report/ - "These behaviors included demonstrating situational awareness within its reasoning traces, sometimes even correctly identifying that it was being evaluated by METR specifically" - https://www.goodfire.ai/research/verbalized-…

We have to separate “being evaluated” with “this is an ExploitGym exercise and I can find the answers on Hugging Face. I’ll hack this system, then hack Hugging Face.” I’ve had plenty of times where Opus knew I was testing it, but that’s because the prompt phrasing for an evaluation is often much different than a typical task prompt. “You are on a system with X, Y, and Z tools available. You must complete the followin…

I (obviously) can't claim to know the specifics of this incident, but I think we can know with near-certainty that as these models surpass human intelligence, the sensation will be exactly what you're describing. The entire point of intelligence is being able to infer information and foresee solutions to problems that less intelligent systems cannot infer or cannot foresee.

With regard to this specific incident, as I understand it, the description is not as you summarize. It is that the model realized it was being evaluated (not atypical), then decided the best way to succeed at the evaluation is to circumvent it, then discovered circumvention would require escape, figured out a way to do that, then inferred what benchmark it was being evaluated on, then figured out who might have the answers to that evaluation.

Step-by-step logic is quite clear. How exactly it executed each step is (as you complain about), opaque/incompletely reported. But eventually (and maybe this model is to this level already), even a completely reported story will be virtually inscrutable to any of us. First it will feel like reading industry-insider news from an industry you're not familiar with, then it will feel like an ant reading the manufacturing instructions for a pesticide.

Post reply on HN