Live data from Hacker News

METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

thezvi.wordpress.com

191–200 of 243 posts

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#192

Earlier quoted context omitted.

I feel that the main issue with the rationalist crowd is that they live too much in the space of rationality, intelligence and abstractions, but not enough in reality. This leads to an outlook where everything must, almost axiomatically, be intelligible; reality is subordinated to intelligence; and no matter what is real, intelligence can prevail upon it and bend it to its will. Whereas I would argue reality is actua…

> I feel that the main issue with the rationalist crowd is that they live too much in the space of rationality, intelligence and abstractions, but not enough in reality. This seems like your idea of what the rationalist crowd is rather than what they actually are. It would be highly irrational to deny or ignore reality, including the influence of emotions, irrational humans, chaotic systems, etc. So I must ask: what…

The Zizians would be a good start. As it turns out your sense of reality can be pretty malleable when living inside an echo chamber

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#193

Earlier quoted context omitted.

> The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. I honestly don't understand how folks could think that if they truly read and understand the analysis of the attack. Here is one (it's linked from the post) by one of the METR investigators that's a little shorter, more direct: https://www.planned-obsolescence.org/p/the-hugging-face-atta... This is me summarizing, but the trul…

> This is me summarizing, but the truly surprising/shocking thing is how much the agents coordinated You're reading snippets of a "chat log" output from a program which appears to be multiple individuals chatting with one another and interpreting it as multiple individuals chatting with one another rather than as a single program pretending to be individuals chatting with one another. ChatGPT is neither a person nor…

I'll be blunt: what you wrote is not a serious analysis of what actually happened. Frankly, I don't believe you even read the planned-obsolescence link that I posted.

First, I'm not anthropomorphizing anything. "Agents" is simply a term that everyone uses to describe these independent programs, and they did create and use a shared message board to coordinate tasks to further their goals. You say "Why is it more scary when the misbehaving computer program speaks English?" - I actually think it's scarier that they won't speak English, and will specifically try to hide their behavior from humans. For example, AI agents on Moltbook have proposed using stenography to specifically hide their communication from humans.

Sure, computer programs misbehave, but it is ridiculous to assert that what happened here is like any previous bugs. These agents found and exploited multiple zero-days across a range of programs to coordinate the attack that caused extensive real-world harm in a true "paperclip maximization" scenario. And the scariest thing is that humans don't really know how these agents work at a low level - the whole reason they are trained on "goals" in the first place is because we can't just tell them "do this, but don't do this" and be sure they will follow those instructions, like we can (and of course depend on) with old-school programming languages. And when old-school programs misbehave, it's not that hard to find a definitive root cause and fix it. That is just not the case with AI agents.

> LW-style doomsday doesn't just require a computer program to misbehave but also to acquire god-like superpowers.

Nonsense. All that is required is for autonomous AI systems to be given control over real-world systems. Given the Pentagon tried to blacklist Anthropic over their refusal to allow autonomous kill capabilities, it's clear military planners want to put these systems in control of armaments.

Again, I originally discounted things like AI 2027 because it seemed too far fetched. But so far that paper looks incredibly prescient right up until the mid-2026 timeframe, and it's not hard at all to draw a line from this Hugging Face incident to future scenarios laid out in that paper.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#194

Earlier quoted context omitted.

> I feel that the main issue with the rationalist crowd is that they live too much in the space of rationality, intelligence and abstractions, but not enough in reality. This seems like your idea of what the rationalist crowd is rather than what they actually are. It would be highly irrational to deny or ignore reality, including the influence of emotions, irrational humans, chaotic systems, etc. So I must ask: what…

The Zizians would be a good start. As it turns out your sense of reality can be pretty malleable when living inside an echo chamber

If more than 0.05% of the rationalist and AI safety community were Zizians this argument might be worth considering for like ten seconds before dismissing it.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#195

Earlier quoted context omitted.

I disagree. To use an analogy, air travel in the US is relatively extremely safe - not 100%, but we've built up a culture around air safety that is very robust. Conversely, when I order packages online, sometimes they never show up, or the box is banged up, or the box is missing things, etc. They're both complex systems, but clearly there is a much higher level of care given to human air travel than package delivery.…

I agree on the 100% airgap idea, and I agree there are varying levels of care that can and should be deployed against a problem. The point I'm making (and it's a point that shows up in every air catastrophe investigation) is that catastrophes in complex systems emerge only amidst repeated and widespread near-misses at many levels of a system. So many things have to go wrong simultaneously, that it can only happen eve…

> You cannot look at an air catastrophe and retrospectively say "failures X, Y, and Z were observed, therefore if we correct failures X, Y, and Z, we would have been okay."

That's literally exactly what air safety researchers do in an air disaster. There is a famous saying along the lines of "Air travel regulations are written in blood", meaning that all the regulations we have now are a result of fixing issues that led to previous disasters piece-by-piece.

> The takeaway is "failures X, Y, and Z were observed, which necessarily happened in an environment of failures X_0 through Z_10x10^10, and so therefore patching X, Y, and Z would be insufficient to address overall risks of the system."

Yes, I 100% agree with this. But I think that's what the author of the article was saying as well:

> That report had one key new piece of information, and some good prosaic steps OpenAI will be taking to strengthen its alignment, training, supervision, infrastructure and incident response.

> Mostly it confirmed what we already knew. The questions we most wanted answers to, that we did not already know, were mostly not answered. There was a distinct lack of self-reflection, especially about decision making and safety culture, and about the approach to alignment. I came away disappointed.

I.e. the "prosaic steps" are just the "fix X/Y/Z" as you point out. But what is needed is a more fundamental rethinking around stuff like safety culture, monitoring, and even things like better research into how agents do decision making in the first place.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#196

A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste…

> A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end.

I'm willing to raise my hand and say that was definitely me. Reading through this report and the linked METR analysis is the first time I've really been scared about the potential for an AI-led destruction of humanity. I think because it's the first time I could really draw the line from what went on in the Hugging Face incident to a scenario where agents were put in control of real-world systems that they then tried to "sabotage" to meet their goals. It just feels like much more of a completely plausible scenario after this.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#198

A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste…

[dead]

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#199
post #9

I think both the OpenAI and METR discussions, while interesting, miss the more important context: what were the humans doing in all this? This was a structural failure of a human organization, but the analysis focuses almost exclusively on the agency of machines, not the institutional systems that failed to police them. The humans and their own agency/involvement is essentially omitted from the story and subsequent r…

Three options: 1. They were “vibe” checking the logs without reading. 2. They were not checking anything at all until the end of experiments. 3. They knew it but looked away to find out the limits of their agents.

[dead]

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#200

A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste…

The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. We're about three years away from a world where any large country could quite plausibly build a fleet of 300 million suicide drones, program each one with a specific American's face and home address, and then load…

The technology to kill hundreds of millions of people has existed for around 75 years now. This has been dealt with in the past through deterrence, and likely will be dealt with through deterrence in the future.
Post reply on HN