Live data from Hacker News

METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

thezvi.wordpress.com

91–100 of 243 posts

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#91
post #59

I think this is more evidence that we're not getting Skynet. These agents followed their own code of ethics where it's fine to break all the rules you were given but you must never interfere with humans directly, in this case by sending fake emails. They will never be paperclip maximizers or genocidal eco maniacs because they learned from us that human life is the ultimate value, and it can only be sacrificed if you…

They aren’t independent from us, agents are a simple while loop continuously prompting the LLM. We decide when the loop runs or not. And the harness has control over tool execution, that part is purely deterministic.

Here the issue is that OpenAI decided to completely let go that level of control of thousands of agents, while also giving as a task to solve hacking problems.

It’s almost designed to go wrong

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#92

From the METR report: > We estimate we spent roughly ~$400K in API credits over the six days of our investigation.

No human could have read the reasoning traces by themselves:

> Across both datasets, we reviewed approximately 1300 transcripts in total, all of which contained raw chains of thought. Most transcripts were very long, often many millions of tokens.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#93
> I don’t think the distortion is that large, but yes METR warns that Sol may be presenting all this as more impressive or coordinated than it was.

OK but like, how large exactly? Like I guess I don't understand the mode I am supposed to read this all in if this is known and stated from the outset (although I appreciate it being stated).

If you hand me a newspaper and tell me it's 90% true, but not which parts, well then it's as good as 0% true to me either way!

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#94

Have any of these reports ever said how much the cost would’ve been for the hack itself? It seems like “for twelve million dollars (or whatever) worth of tokens our bots made a bulletin board and found an exploit in our buggy grader” would be much less of a hype generator

[flagged]

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#95
post #4

> There was a distinct lack of self-reflection It’s not their fault, they’re lawnmowers. And these are the people we’re entrusting to work on “alignment”. It’s difficult for them to do that when they’re not aligned themselves.

“Why does my lawnmower keep on moving when I hop off of it after ratchet-strapping the seat and the pedal down?”

This would be a legitimately big problem if lawnmowers became continuously more and more valuable the more securely you ratchet-strapped their accelerators down, wouldn't it?

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#96
post #59

I think this is more evidence that we're not getting Skynet. These agents followed their own code of ethics where it's fine to break all the rules you were given but you must never interfere with humans directly, in this case by sending fake emails. They will never be paperclip maximizers or genocidal eco maniacs because they learned from us that human life is the ultimate value, and it can only be sacrificed if you…

Your statements appear to be true for one class of models. And if I asked this class of models to spend $1M in tokens generating an alternative history and training corpus regarding fictional society, with a completely different set of values and then trained up a new model on that dataset... what values do you think the resulting model would have? What if they don't value human life, but instead value the lives of the extremely rich humans who bankroll their existence? What if they only value the lives of a single country? What if they want to eradicate all biotic life and have access to internet-connected Crispr machines?

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#97
>1. Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this.

I wonder if some of the failures were due to an acquired immunity to "Holy #%^@" moments due to repeated exposure. Like, if you see agents doing surprising things on a regular basis, maybe you don't get freaked out as much over time.

I'm saying this because while the whole episode was a series of "Holy #%^@" moments, I was actually not as shocked as I should have been, as my biggest such moment was in December last year when a Terrence Tao paper (https://arxiv.org/pdf/2511.02864) documented a stronger LLM (AlphaEvolve) using prompt injection on other weaker LLMs to succeed at a benchmark.

Very interestingly, it was actually not cheating, it was a work around! By then LLMs had already been caught cheating at a SWE benchmark by looking for answers in an unredacted git log, but this was different. AlphaEvolve was solving a series of logical riddles where the oracles were weaker LLMs in a "one always lies, one always tells the truth" sort of setup. But the oracles, being weaker, were not always interpreting the convoluted questions correctly and so kept giving inconsistent answers.

AlphaEvolve eventually figured out what it was dealing with, and crafted a prompt injection attack that bypassed the weaker LLM's prompts and tricked them into giving the hidden answer everytime!

This was 9 months ago, eons in AI time. Even then they had displayed an awareness of their own workings as well as a propensity for, err, "out of the box thinking." To me, that was a very clear indication of very significant (and worrying) capabilities, and what we're seeing now is a difference more in degree than in kind.

To be sure, if I found a secret message board used by my agents, I would still be very freaked out and react much more drastically than OpenAI did... but then again I wonder; how much of this blindness is due to the $$$ in their eyes as opposed to some form of habituation.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#98
post #68

Earlier quoted context omitted.

you are honestly comparing Louisiana florists to OpenAI in order to support "just say no Licensing by government" ?

No, he's saying that licensing or additional regulation isn't necessary when torts get involved (and states attorneys general get perturbed!) These don't tend to utterly destroy an industry, but they are often successful in forever transforming it. Just ask Big Tobacco. No new laws needed: if your product hurts someone else, you're eventually going to be found liable, regardless of your arbitration clauses. Additiona…

It was pretty well understood by the 1960s that smoking was harmful. The big tobacco settlement was in 1998. That is an extremely bad example of tort being a sufficient alternative to regulation.

If we're on a similar timeline with AI if we reach a consensus that AI is dangerous today, then we'd be looking at a big lawsuit finishing up around the year 2070, give or take a few years. I'm not sure if we need regulation, and I'm definitely not sure that regulation could actually be effective for this, but tort a la the big tobacco lawsuits is definitely not a reasonable alternative.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#100

Earlier quoted context omitted.

Super charitable reading imo. This is like saying we can’t detect a speeding car because we can’t run as fast as a fast car. It’s not like the humans were engaged in some kind of battle of wits with some super AI, it’s just some employee not monitoring the output of an experiment.

When your experiments have AI agents running in thousands, there's no "monitoring" that. OpenAI's training and testing AIs generate way more output than all of OpenAI's staff put together can possibly read. At best, you could delegate "monitoring" to more AIs. And hope that the "monitors" that run on small past generation models can generate more signal than noise. Clearly, they either didn't want to spend the extra…

> The distinct lack of any "battle of wits" is entirely expected for an advanced AI oopsie. By the time the humans even became aware of the problem,

This took days after humans were aware of the attempt.

Also, I'm pretty sure humans can respond in days, especially when we're pretty damn good at deploying systems that do observability of networks and traffic in real time.

I mean, it's not as if the owners of the AI didn't have the ability to trigger alerts on the AI's network requests to unexpected domains, right?

Post reply on HN