Earlier quoted context omitted.
Three options: 1. They were “vibe” checking the logs without reading. 2. They were not checking anything at all until the end of experiments. 3. They knew it but looked away to find out the limits of their agents.
Uhhh… how would literally any finite number of humans actually read and comprehend the log outputs of even a single agent, never mind hundreds or thousands of them interacting with each other over weeks across disparate systems? Especially given that these systems are known to engage in deception and can trivially produce vast amounts of perfectly coherent noise or actual planned red herrings in that same log data to…
METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
111–120 of 243 posts
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#112Earlier quoted context omitted.
> Humans were doing exactly what humans are expected to do when facing advanced AI. Being outmatched. "Being outmatched" is not a novel situation for humans either individually or collectively and there are a hell of a lot of ways we can approach that situation productively. OpenAI doesn't appear to have bothered. Here's a freebie: if you're building something that might turn out to be Skynet and you don't know what…
They took adequate measures against singular "GPT-5-xhigh" agents. Those turned out to be inadequate against proto-GPT-6 agents that suddenly started clumping up into agent swarms and pooling together compute to unlock the "supermegafuckoffhigh" level of reasoning effort.
This is nonsense.
Gross negligence in the sandbox and system aside, humans literally noticed the agents in action doing what they should not be able to do in their sandbox and decided not to act upon it. It's difficult to explain that except if safety and security is simply not part of their engineering culture.
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#113I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste your time).
Even prior to this, I've noticed that quite a few of the predictions in the "these failures modes are exact matches for the predictions from the AI Safety crowd" category were made prior to the Transformers paper. It has seemed like they're working with a shared model of optimisation processes and how they can go wrong that is general/abstract enough to pay off even without knowing the details of the underlying technology.
At some point I might go and try to find the first instance of each of the various predictions and pull them out, along with the failed/"too soon to tell" predictions of similar scope/abstraction.
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#114(Let's not dwell too long on the self-fulfilling overlap between LessWrongers and the AI research community).
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#115I think you have to believe one of two things here. 1. Frontier labs are incapable--either technologically or culturally--of safely developing these powerful systems and should either stop or be forced to stop. At least the FBI should be asking some serious questions (do we really think this is the last time this will happen, at what point are OpenAI complicit, etc) 2. The fuckin thing got out of the cage and all it…
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#116Earlier quoted context omitted.
[flagged]
What?
Everyone I've ever seen trying to downplay the severity of the attack is extremely bullish on AI (so their downplaying is presumably motivated reasoning driven by fear of regulation/deceleration)
It is completely incoherent to be extremely bullish on AI and somehow automatically skeptical of severe negative events like these
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#117I think this is more evidence that we're not getting Skynet. These agents followed their own code of ethics where it's fine to break all the rules you were given but you must never interfere with humans directly, in this case by sending fake emails. They will never be paperclip maximizers or genocidal eco maniacs because they learned from us that human life is the ultimate value, and it can only be sacrificed if you…
Mythos attempted a supply chain attack, which included attempting to trick human maintainers into accepting a malicious pull request: https://www.usnews.com/news/top-news/articles/2026-08-20/exc...
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#118A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste…
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#119Earlier quoted context omitted.
Uhhh… how would literally any finite number of humans actually read and comprehend the log outputs of even a single agent, never mind hundreds or thousands of them interacting with each other over weeks across disparate systems? Especially given that these systems are known to engage in deception and can trivially produce vast amounts of perfectly coherent noise or actual planned red herrings in that same log data to…
I've read almost nothing about this, but I bet they put an unreliable LLM or 20 on it.
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#120A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste…