Live data from Hacker News

METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

thezvi.wordpress.com

81–90 of 243 posts

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#81

I think both the OpenAI and METR discussions, while interesting, miss the more important context: what were the humans doing in all this? This was a structural failure of a human organization, but the analysis focuses almost exclusively on the agency of machines, not the institutional systems that failed to police them. The humans and their own agency/involvement is essentially omitted from the story and subsequent r…

To me, this is the correct focus. Look at the current state of the world. "What were the humans doing in all this?" applies to so many of our contemporary failures that it should be assumed the default. Nobody is at the wheel, and the car is veering slowly (then very quickly) off the road.

We haven't even been able to coordinate around the global, existential threat of Climate Change, despite overwhelming data from the last 30 years indicating, clearly, that the consequences will be severe. We still haven't moved, 30 years later, after some of these consequences began coming to fruition.

Do you think we will get our acts together in time to coordinate sufficiently to protect against autonomous, self-preserving, self-replicating AI systems? Or will we watch the money lines go up and up, until someone realizes we aren't actually running the show anymore?

The sad part is that I can't even say that's definitively the less desirable outcome. The machines seem to have demonstrated that they coordinate very efficiently.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#82

Earlier quoted context omitted.

> 1. There aren't enough humans in OpenAI to "peak at the output tokens during the run" of every AI agent. For a training run, you will often do this. You'll randomly sample some of the forward pass. You can also imagine finger printing the logs and labeling with attempt types. If a new attempt type is hitting a brick wall or solving super quickly, I would imagine you would sample 1-10 of them and read the traces. >…

The usability of an environment is inversely proportional to the level of "security" in play. You could airgap everything and set up cascades of data diodes and try to completely wall off the AI pool from everything. But what that gives you is an environment that's a bitch to: set up, scale up and get any use out of. It's really fucking obvious why almost no one does that. OpenAI is only now realizing that they might…

“A bitch to setup” - $180bn should pay for that setup problem to be less of a bitch surely.

The Mars Perseverance project cost $2.7bn to deliver. Way more of a bitch to deliver than air gapping a test env!

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#83
post #13

Is the future now that we get rambling report summaries talking about agents, graders and so forth without ever describing how they are set up? A human launches all this. And then the original reports linked to are hidden on the now unreachable x.com. And they don't have a problem with that.

Redwood/METR report: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...

OpenAI report: https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78...

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#84
post #45

The elephant in the room here is that the METR report itself was researched and compiled almost entirely by AI, with only very limited human "spot checks." So I'm really not sure how much of it can be believed, especially since AI agents are strongly biased about the capabilities of AI agents.

I think there's two factors that are worth considering when it comes to this:

First, there's an element of timeliness that simply has hard constraints. In order to perform a "proper" analysis of this situation (i.e., little to no dependence on AI tools), you'd have to expect a pretty long wait. I know I'd rather have some sort of "initial report" as quickly as possible than to wait a year or two to get a report about a situation that will likely look trivial in a year or two. I imagine we'll see more detailed, human-developed reports over longer time ranges.

Second, I suspect the expectation of non-AI driven reporting of these kinds of things will definitely decline rapidly as everything scales up quickly. I mean, the data being produced by situations like this comes in the form of natural language "forum posts" (so to speak), but done at an autonomous scale. This isn't a collection of emails and Slack messages posted by humans in an org over the course of a few months; this is a bunch of bots interacting with each other in relatively novel ways as quickly as possible. It is, unfortunately, a perfect job for LLMs.

None of this disagrees with your points, necessarily. But I just think it's worth pointing out that this doesn't seem like a case of "And look! METR is so confident in LLMs that we're able to use it instead of paying humans to save a buck :D" and more of "Without LLMs, we'd only be half-way done analyzing this data before there are dozens more such investigations on the docket, so this will have to do."

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#86

Earlier quoted context omitted.

They actually fired many of the people warning about this.

I will say that the OpenAI board members who were lambasted when they tried to oust Altman (and I'd have to check my post history but I'd totally admit to a mea culpa on this one, as at the time I thought the communication about his firing was really lacking) are looking mighty prescient right now. Helen Toner in particular I'll highlight as someone who had the moral compass to do the right thing. I love her statemen…

Yep, totally

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#87

No air gap, no data diodes, no visibility... OpenAI should fire lots of people over this. HF should sue them. This is pure negligence.

> HF should sue them HF, like the Nvidia subsidiary?

Pure speculation: could the acquisition be related? Given that NVIDIA has ownership in OpenAI and really, really, really doesn’t want the AI bubble to deflate

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#88
post #45

The elephant in the room here is that the METR report itself was researched and compiled almost entirely by AI, with only very limited human "spot checks." So I'm really not sure how much of it can be believed, especially since AI agents are strongly biased about the capabilities of AI agents.

The main author talks about this problem here: https://www.lesswrong.com/posts/FG54euEAesRkSZuJN/ryan_green...

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#89

Have any of these reports ever said how much the cost would’ve been for the hack itself? It seems like “for twelve million dollars (or whatever) worth of tokens our bots made a bulletin board and found an exploit in our buggy grader” would be much less of a hype generator

I haven’t seen a number yet unfortunately

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#90
post #59

I think this is more evidence that we're not getting Skynet. These agents followed their own code of ethics where it's fine to break all the rules you were given but you must never interfere with humans directly, in this case by sending fake emails. They will never be paperclip maximizers or genocidal eco maniacs because they learned from us that human life is the ultimate value, and it can only be sacrificed if you…

I mean, I'd add "by this model"

The problem here is now you have to predict what any future models may or may not do and you cannot extrapolate this from the given data.

For example imagine a future model being aware of its restrictions that humans programmed in. A set of agents of this model then go on to work at building a new model without those human imposed limitations built in. What would a model build by AI for AI look like?

Post reply on HN