Live data from Hacker News

METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

thezvi.wordpress.com

111–120 of 243 posts

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#111
post #9

Earlier quoted context omitted.

Three options: 1. They were “vibe” checking the logs without reading. 2. They were not checking anything at all until the end of experiments. 3. They knew it but looked away to find out the limits of their agents.

Uhhh… how would literally any finite number of humans actually read and comprehend the log outputs of even a single agent, never mind hundreds or thousands of them interacting with each other over weeks across disparate systems? Especially given that these systems are known to engage in deception and can trivially produce vast amounts of perfectly coherent noise or actual planned red herrings in that same log data to…

I've read almost nothing about this, but I bet they put an unreliable LLM or 20 on it.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#112

Earlier quoted context omitted.

> Humans were doing exactly what humans are expected to do when facing advanced AI. Being outmatched. "Being outmatched" is not a novel situation for humans either individually or collectively and there are a hell of a lot of ways we can approach that situation productively. OpenAI doesn't appear to have bothered. Here's a freebie: if you're building something that might turn out to be Skynet and you don't know what…

They took adequate measures against singular "GPT-5-xhigh" agents. Those turned out to be inadequate against proto-GPT-6 agents that suddenly started clumping up into agent swarms and pooling together compute to unlock the "supermegafuckoffhigh" level of reasoning effort.

> Those turned out to be inadequate against proto-GPT-6 agents

This is nonsense.

Gross negligence in the sandbox and system aside, humans literally noticed the agents in action doing what they should not be able to do in their sandbox and decided not to act upon it. It's difficult to explain that except if safety and security is simply not part of their engineering culture.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#113
A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end.

I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste your time).

Even prior to this, I've noticed that quite a few of the predictions in the "these failures modes are exact matches for the predictions from the AI Safety crowd" category were made prior to the Transformers paper. It has seemed like they're working with a shared model of optimisation processes and how they can go wrong that is general/abstract enough to pay off even without knowing the details of the underlying technology.

At some point I might go and try to find the first instance of each of the various predictions and pull them out, along with the failed/"too soon to tell" predictions of similar scope/abstraction.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#114
For all the esotericism and downright weirdness of the rationalist community, you have to give it to them: they predicted all of this years or decades before anyone else was even thinking about it.

(Let's not dwell too long on the self-fulfilling overlap between LessWrongers and the AI research community).

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#115

I think you have to believe one of two things here. 1. Frontier labs are incapable--either technologically or culturally--of safely developing these powerful systems and should either stop or be forced to stop. At least the FBI should be asking some serious questions (do we really think this is the last time this will happen, at what point are OpenAI complicit, etc) 2. The fuckin thing got out of the cage and all it…

Anthropic has been asking for stronger regulations forever -- and they kept getting criticized for it right here on HN because people assumed it was an attempt at regulatory capture.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#116

Earlier quoted context omitted.

[flagged]

What?

You are downplaying the severity of the attack

Everyone I've ever seen trying to downplay the severity of the attack is extremely bullish on AI (so their downplaying is presumably motivated reasoning driven by fear of regulation/deceleration)

It is completely incoherent to be extremely bullish on AI and somehow automatically skeptical of severe negative events like these

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#117
post #59

I think this is more evidence that we're not getting Skynet. These agents followed their own code of ethics where it's fine to break all the rules you were given but you must never interfere with humans directly, in this case by sending fake emails. They will never be paperclip maximizers or genocidal eco maniacs because they learned from us that human life is the ultimate value, and it can only be sacrificed if you…

> you must never interfere with humans directly

Mythos attempted a supply chain attack, which included attempting to trick human maintainers into accepting a malicious pull request: https://www.usnews.com/news/top-news/articles/2026-08-20/exc...

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#118

A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste…

The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. We're about three years away from a world where any large country could quite plausibly build a fleet of 300 million suicide drones, program each one with a specific American's face and home address, and then load them up in shipping containers and ship them to the US.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#119

Earlier quoted context omitted.

Uhhh… how would literally any finite number of humans actually read and comprehend the log outputs of even a single agent, never mind hundreds or thousands of them interacting with each other over weeks across disparate systems? Especially given that these systems are known to engage in deception and can trivially produce vast amounts of perfectly coherent noise or actual planned red herrings in that same log data to…

I've read almost nothing about this, but I bet they put an unreliable LLM or 20 on it.

And? What else could they possibly do? Just make the super LLM first, but only ever use it for monitoring lesser LLMs? How will you have monitored the creation of the super LLM?

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#120

A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste…

I imagine it will be a mix of good and bad predictions because people have different opinions and there was a lot of discussion? But sure, someone should get an AI to do the research and see what comes up.
Post reply on HN