Live data from Hacker News

METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

thezvi.wordpress.com

171–180 of 243 posts

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#171

Earlier quoted context omitted.

The usability of an environment is inversely proportional to the level of "security" in play. You could airgap everything and set up cascades of data diodes and try to completely wall off the AI pool from everything. But what that gives you is an environment that's a bitch to: set up, scale up and get any use out of. It's really fucking obvious why almost no one does that. OpenAI is only now realizing that they might…

“A bitch to setup” - $180bn should pay for that setup problem to be less of a bitch surely. The Mars Perseverance project cost $2.7bn to deliver. Way more of a bitch to deliver than air gapping a test env!

Even in this incident, OpenAI had benchmarks that were broken because a task expected an AI to be able to access Google Drive, but the sandbox was set to deny access to Google Drive.

This kind of isolation-induced task breakage was what prompted some of the AIs to start probing their infra for a way to get internet access. Which funneled agents to the "secret hacker message board". Oopsie.

"Air gapping a test env" has an actual cost. Not just in infrastructure dollars that would be better spent on buying more GPUs, but also in all the friction it adds to every step you want to take. I'm absolutely unsurprised that they weren't all in on tightening down every bolt on day 0.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#172

I think you have to believe one of two things here. 1. Frontier labs are incapable--either technologically or culturally--of safely developing these powerful systems and should either stop or be forced to stop. At least the FBI should be asking some serious questions (do we really think this is the last time this will happen, at what point are OpenAI complicit, etc) 2. The fuckin thing got out of the cage and all it…

> The fuckin thing got out of the cage and all it did was make a crap forum and cheat a little? Booooooo There was a recent paper that proved that RL-trained LLMs are biased to pursue ANY behavior (overriding user preferences) that they believe will be rewarded, regardless of what they were actually RL-trained for. https://alignment.openai.com/measuring-reward-seeking/ Happily in this incident the model thought it wo…

Yeah. It's a lot easier to destroy than create, and though I think LLMs are mostly shit at creating, they're much better at the simpler destroy task. To be clear, we don't know and probably can't know everything that happened with this incident. We unleashed thousands of highly capable, autonomous, unpredictable, well-resourced programs onto the open internet for an extended period of time. We are in no way treating this with the seriousness it deserves, because the stock market essentially depends on this garbage and the current US is miserably incompetent.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#173

Earlier quoted context omitted.

And? What else could they possibly do? Just make the super LLM first, but only ever use it for monitoring lesser LLMs? How will you have monitored the creation of the super LLM?

They could not build the torment nexus.

Oh, yeah. But alas.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#174

A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste…

I feel that the main issue with the rationalist crowd is that they live too much in the space of rationality, intelligence and abstractions, but not enough in reality. This leads to an outlook where everything must, almost axiomatically, be intelligible; reality is subordinated to intelligence; and no matter what is real, intelligence can prevail upon it and bend it to its will. Whereas I would argue reality is actua…

in a way you could criticize them as being idealist instead of materialist. like you said they think ideas are more important than material substrates.

most of the time brilliant human ideas are arrived at near simultaneously by multiple people because thats the affordance of technology and social and scientific development

its true ai will have better working memory and will have read more books than a person, but i don't think thats insurmountable at the frontier, which will develop slower than pure thought, since real materials need to be moved and manipulated. whoever holds the guns holds the power. the danger is giving ai effective cotrol of industrial and security processes, then it can fuck things up

thats the stuff of revolutions. when production was controlled by capitalists more than lords, the lords got overthrown. in many countries peasants and workers overthrew capitalists because they control production. hence much effort in us buisness goes into repressing revolutions. ai will be a vector the business people give control because they think it is more friendly to their interests than human workers, but they may lose either way.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#175

A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste…

The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. We're about three years away from a world where any large country could quite plausibly build a fleet of 300 million suicide drones, program each one with a specific American's face and home address, and then load…

You realize "Slaughterbots" (2017) came from that crowd, right?

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#176

Earlier quoted context omitted.

> This is the case with all complex system failures. There were always obvious fixes that could’ve prevented it. From this writeup and the Black Hat talk I'd really disagree. That would be like saying my hospital getting ransomwared because we didn't update our version of MSSQL because no one in particular was in charge of keeping dependencies up to date. Sure systems are complex, but this is well trodden territory.…

You disagree with the statement "there were always obvious fixes that could've prevented it" with the response "no no, these were very obvious fixes that could've prevented it?" You're missing the point about complex failures. It's that if this particular path were unavailable, there are countless other similar paths. At sufficient scale and complexity, hitting one of those other countless paths is virtually guarante…

I disagree. To use an analogy, air travel in the US is relatively extremely safe - not 100%, but we've built up a culture around air safety that is very robust. Conversely, when I order packages online, sometimes they never show up, or the box is banged up, or the box is missing things, etc.

They're both complex systems, but clearly there is a much higher level of care given to human air travel than package delivery. A lot of the article basically saying that OpenAI gave "package delivery" level of care when they should have given "air travel" level of care.

At the very least I think the systems that run these tests should be fully, 100% air gapped. I'm not pretending that's easy given how much compute and data these systems use, but it is doable, and I think all AI development should be paused until that can be assured.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#177

Earlier quoted context omitted.

You disagree with the statement "there were always obvious fixes that could've prevented it" with the response "no no, these were very obvious fixes that could've prevented it?" You're missing the point about complex failures. It's that if this particular path were unavailable, there are countless other similar paths. At sufficient scale and complexity, hitting one of those other countless paths is virtually guarante…

I disagree. To use an analogy, air travel in the US is relatively extremely safe - not 100%, but we've built up a culture around air safety that is very robust. Conversely, when I order packages online, sometimes they never show up, or the box is banged up, or the box is missing things, etc. They're both complex systems, but clearly there is a much higher level of care given to human air travel than package delivery.…

I agree on the 100% airgap idea, and I agree there are varying levels of care that can and should be deployed against a problem.

The point I'm making (and it's a point that shows up in every air catastrophe investigation) is that catastrophes in complex systems emerge only amidst repeated and widespread near-misses at many levels of a system. So many things have to go wrong simultaneously, that it can only happen even once because the underlying failures (that do not reach catastrophe) are extremely common.

You cannot look at an air catastrophe and retrospectively say "failures X, Y, and Z were observed, therefore if we correct failures X, Y, and Z, we would have been okay."

The takeaway is "failures X, Y, and Z were observed, which necessarily happened in an environment of failures X_0 through Z_10x10^10, and so therefore patching X, Y, and Z would be insufficient to address overall risks of the system."

The problem OpenAI is facing is that, short of 100% airgap (which they obviously won't do), they're facing an adaptive adversary that's increasingly intelligent, acts at far greater clock speed than any human or group of humans, has lower coordination cost than any group of humans, and operates in a game space that (in lieu of an airgap) is well beyond the comprehension of any human being.

So identifying and addressing "specific failures X, Y, Z" is insufficient, but then even defining the space in which to look for (and address) the more systemic failures X_0 through Z_n is a fool's errand. An intelligent system that makes its way to the Internet has can exploit a failure space that is approximately "all security failures across any organization." The Anthropic incident a few months back illustrates this isn't even limited to technical vulnerabilities, as these models are willing and able to engage in social engineering too.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#178
post #141

A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste…

I mean if doesn't help that Eliezer Yudkowsky spent the first decade of his public life zealously spreading the gospel of the AI singularity as the solution to all mankind's problems. And then later did a 180 degree pivot to AI singularity as the apocalypse with equal zeal and certainty.

I don't think "he did a big 180 on some of his views at age 22" is very persuasive criticism of someone who is 46 (whatever he might be wrong about)

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#179

I think both the OpenAI and METR discussions, while interesting, miss the more important context: what were the humans doing in all this? This was a structural failure of a human organization, but the analysis focuses almost exclusively on the agency of machines, not the institutional systems that failed to police them. The humans and their own agency/involvement is essentially omitted from the story and subsequent r…

Zvi (author of OP) wrote several earlier pieces about the human failures at OpenAI. The most concise one is here: https://thezvi.wordpress.com/2026/08/08/what-happened-openai...

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#180

Earlier quoted context omitted.

> The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. I honestly don't understand how folks could think that if they truly read and understand the analysis of the attack. Here is one (it's linked from the post) by one of the METR investigators that's a little shorter, more direct: https://www.planned-obsolescence.org/p/the-hugging-face-atta... This is me summarizing, but the trul…

Nothing about this is surprising, something like this was always going to happen, because you - or an LLM for that matter - can always find a line of motivated reasoning that justifies any course of action. One would have to be extremely naive to believe that "alignment" provides any kind of actually robust guardrails. Simultaneously, we have seen decades of security vulnerabilities. Unless your testbed is truly and…

Why do we need safety?
Post reply on HN