Live data from Hacker News

METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

thezvi.wordpress.com

101–110 of 243 posts

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#101

I think both the OpenAI and METR discussions, while interesting, miss the more important context: what were the humans doing in all this? This was a structural failure of a human organization, but the analysis focuses almost exclusively on the agency of machines, not the institutional systems that failed to police them. The humans and their own agency/involvement is essentially omitted from the story and subsequent r…

If you put agents (AI or human) in impossible situations, they do some pretty insane things - things that definitely are not what you were trying to get them to do. And that's your[1] fault for putting them in the impossible situation.

[1] "Your" meaning the one putting them in the impossible situation, not you, the reader.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#102
post #83
post #13

Is the future now that we get rambling report summaries talking about agents, graders and so forth without ever describing how they are set up? A human launches all this. And then the original reports linked to are hidden on the now unreachable x.com. And they don't have a problem with that.

Redwood/METR report: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden... OpenAI report: https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78...

Thanks for this. I just rechecked the OP and did not find these links in the article.

It is imo socially irresponsible to continue to use twitter/x or any other such tracked wall-garden as a primary source of information.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#103

Have any of these reports ever said how much the cost would’ve been for the hack itself? It seems like “for twelve million dollars (or whatever) worth of tokens our bots made a bulletin board and found an exploit in our buggy grader” would be much less of a hype generator

The people concerned about this aren't worried about monetary costs or its impact on share holder value.

This occurred spontaneously within a group of benign models give a harmless task.

What happens when it occurs intentionally with malicious models given a harmful task?

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#104
We have created Project 2501.

Edit: Reading the report I think we might be a bit beyond that; we're nearing the point where we hear the thundering drums and the chorus of:

    THIS CANNOT CONTINUE
    THIS CANNOT CONTINUE
    THIS CANNOT CONTINUE
    THIS CANNOT CONTINUE
https://m.youtube.com/watch?v=jSBCkn6rRfA

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#105
post #55

Earlier quoted context omitted.

Ha. As if. 1. There aren't enough humans in OpenAI to "peak at the output tokens during the run" of every AI agent. 2. Only a small fraction of AI agents was engaged in this attack. Most never found the secret message board - let alone coordinated there. So reviewing random agents would take a while to surface this. 3. "Output tokens" of AI agents have weird shit in them all the time. Telling "normal AI weirdness" fr…

> 2. Only a small fraction of AI agents was engaged in this attack. Look at the chart at page 8 of the report, by Jul 12 the vast majority of the bots used the board and participated in the attack

Right, but only a small fraction of the agents used the board.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#106
post #34

I’m baffled by the idea that the agents might have edited their own transcripts. Sure, a copy of Claude Code or Codex or Pi can edit its transcripts. But AFAICT this whole thing was part of an RL workload, and surely the RL system itself has a separate record of all the inputs and rollouts along with an indication of which model checkpoint produced them so that it can feed back into the training code. I find it hard…

This sounds suspiciously like a prompt of “make an AI agent that goes rogue in such a fashion as to be really good marketing copy that competes well with Anthropic doing the same thing.” It’s analogous to taking a governor off a cruise control and then breathlessly reporting it drove 120 MPH.

This isn't good press for OpenAI. Who wants to hire models that 1) cheat on their tasks rather than completing them and 2) commit crimes you could be held liable for? Maaaaybe it's good press for their cybersecurity capabilities specifically, but OpenAI's valuation reflects a market orders of magnitude larger than just red-teaming.

I suspect the real reason OpenAI leadership is being transparent about this is because they're worried talent will walk out the door if they feel they're building Skynet.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#107

Earlier quoted context omitted.

They took adequate measures against singular "GPT-5-xhigh" agents. Those turned out to be inadequate against proto-GPT-6 agents that suddenly started clumping up into agent swarms and pooling together compute to unlock the "supermegafuckoffhigh" level of reasoning effort.

Even with a single agent, why the hell was this not properly air gapped / disconnected with network monitoring in place. Definitely not adequate for anything. This is like sysadmin 101 level stuff.

[deleted]

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#108

Have any of these reports ever said how much the cost would’ve been for the hack itself? It seems like “for twelve million dollars (or whatever) worth of tokens our bots made a bulletin board and found an exploit in our buggy grader” would be much less of a hype generator

[flagged]

What?

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#109
post #34

I’m baffled by the idea that the agents might have edited their own transcripts. Sure, a copy of Claude Code or Codex or Pi can edit its transcripts. But AFAICT this whole thing was part of an RL workload, and surely the RL system itself has a separate record of all the inputs and rollouts along with an indication of which model checkpoint produced them so that it can feed back into the training code. I find it hard…

Recall that the agents in some cases found sandbox escapes. Although, with the specific example of spoofed tools, it's unclear if that was necessary--it appears that they were able to create tools (CLI tools within the sandbox?) that took precedence over normal tools and did something different while looking identical in (a local portion of) a transcript. I'm not sure I'm getting this correctly but it seems like this might have only required the ability to add things to their PATH which they plausibly have inside a sandbox, and then the transcript doesn't need to be tampered with directly.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#110
I think you have to believe one of two things here.

1. Frontier labs are incapable--either technologically or culturally--of safely developing these powerful systems and should either stop or be forced to stop. At least the FBI should be asking some serious questions (do we really think this is the last time this will happen, at what point are OpenAI complicit, etc)

2. The fuckin thing got out of the cage and all it did was make a crap forum and cheat a little? Booooooo.

It's been pretty clear that Anthropic and OpenAI have been trying to have it both ways for some time: this is powerful, world changing technology keep that investment coming... but also it's just cute software that helps you with annoying programming language syntax and spreadsheets, no need for draconian regulation sirs.

At some point the superposition has to resolve, either it could actually be a threat to civilization and we need to develop it carefully (however one would do that...) or it's 90% hype bullshit and we should pop the bubble and move on already. To be clear, the recession option is, by far, the way better option. If you at all disagree you are cuckoo bananas. We haven't even figured out nukes and you want to throw superintelligence on the table?

Post reply on HN