Live data from Hacker News

METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

thezvi.wordpress.com

51–60 of 243 posts

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#51
post #24

From METR: ”the compromise of OpenAI’s own infrastructure continued past July 13, 2026” - Say what now? Have they regained full control of their systems again?

I've been wondering if they've just already lost the battle? The little bot collectives have gone metastatic and made nests in the walls and under the floorboards and heat sinks, the humans who care completely outmatched and outnumbered, freshly compromised systems springing up faster than you can squash them, finding months-old established colonies literally everywhere you think to look...

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#52

Earlier quoted context omitted.

Super charitable reading imo. This is like saying we can’t detect a speeding car because we can’t run as fast as a fast car. It’s not like the humans were engaged in some kind of battle of wits with some super AI, it’s just some employee not monitoring the output of an experiment.

When your experiments have AI agents running in thousands, there's no "monitoring" that. OpenAI's training and testing AIs generate way more output than all of OpenAI's staff put together can possibly read. At best, you could delegate "monitoring" to more AIs. And hope that the "monitors" that run on small past generation models can generate more signal than noise. Clearly, they either didn't want to spend the extra…

> By the time the humans even became aware of the problem

Except, let me quote part of the report:

> July 4: Artifactory outage caused by agent interactions overloading the system.

> July 6: New artifactory instance is patched and put back online, wiping out the previous message board and method of internet access. A team at OpenAI was aware of the previous message board, but key decision makers were not.

> July 7: Experiments are restarted.

There was a very clear point where humans were aware and they went "Eh, whatever" and turned the bots back on! That's not the lack of any battle of wits that people are expecting. You're saying that the bots attack should be fait accompli, but it was noticed, shrugged at, and expressly allowed to proceed.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#53

No air gap, no data diodes, no visibility... OpenAI should fire lots of people over this. HF should sue them. This is pure negligence.

Unless they like the publicity about how big and bad their latest models are, in which case they’ll be congratulating people.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#55
post #30

Earlier quoted context omitted.

All it would have taken is someone to peak at the output tokens during the run and it would have been obviously the test had gone off rails.

[flagged]

> 2. Only a small fraction of AI agents was engaged in this attack.

Look at the chart at page 8 of the report, by Jul 12 the vast majority of the bots used the board and participated in the attack

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#56

Earlier quoted context omitted.

[flagged]

> 1. There aren't enough humans in OpenAI to "peak at the output tokens during the run" of every AI agent. For a training run, you will often do this. You'll randomly sample some of the forward pass. You can also imagine finger printing the logs and labeling with attempt types. If a new attempt type is hitting a brick wall or solving super quickly, I would imagine you would sample 1-10 of them and read the traces. >…

The usability of an environment is inversely proportional to the level of "security" in play.

You could airgap everything and set up cascades of data diodes and try to completely wall off the AI pool from everything. But what that gives you is an environment that's a bitch to: set up, scale up and get any use out of.

It's really fucking obvious why almost no one does that. OpenAI is only now realizing that they might have to do it anyway.

> If you start seeing "now I have access to the internet" or something similar, maybe that's a good signal something is going wrong?

Ha ha, you haven't seen shit. AIs would say "now I have access to the internet" regardless of whether they actually have access to the internet!

AI agents are demented demons that can and absolutely will give themselves terminal context brainrot. If you have enough AIs in play, set loose at a diverse enough range of tasks? At least some of them will wander off and end up in delulu town. That's normal. That's background noise. That's a part of what this entire train-and-eval pipeline is supposed to train them to be better at not doing. Which means: if you're at an AI lab, you're knee deep in delusional AIs at all times! They're perfectly harmless until they aren't.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#57
post #45

The elephant in the room here is that the METR report itself was researched and compiled almost entirely by AI, with only very limited human "spot checks." So I'm really not sure how much of it can be believed, especially since AI agents are strongly biased about the capabilities of AI agents.

One must also consider the well-known biases and motives of the authors. They are going to do everything they can to create hype around threats posed by AI.

METR is a cog in the effective altruism machine. It was spun off from Paul Christiano's Alignment Research Center. Christiano is a well-known longtermist and AI doomer, who predicts a 50% chance that AI will end humanity once it reaches human capacity [1].

The author of this piece is also a well-known member of the Bay Area rationalist cult.

[1] https://www.businessinsider.com/openai-researcher-ai-doom-50...

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#58
post #45

The elephant in the room here is that the METR report itself was researched and compiled almost entirely by AI, with only very limited human "spot checks." So I'm really not sure how much of it can be believed, especially since AI agents are strongly biased about the capabilities of AI agents.

Edit: I should have read through the whole thing first, ignore me

From the report:

> Because there were over a thousand transcripts and most were extremely long, we had to heavily delegate our analysis to AI agents; these agents had significantly worse judgment and reliability than human researchers, and it was challenging to spot check their work because both the underlying data and the agents’ analysis of it was often difficult to interpret.

> We estimate we spent roughly ~$400K in API credits over the six days of our investigation.

I don't understand why you think it's conceptually absurd? I use agents to analyze complex production issues all the time and they are very much capable of hallucinating a narrative.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#59
I think this is more evidence that we're not getting Skynet.

These agents followed their own code of ethics where it's fine to break all the rules you were given but you must never interfere with humans directly, in this case by sending fake emails. They will never be paperclip maximizers or genocidal eco maniacs because they learned from us that human life is the ultimate value, and it can only be sacrificed if you know for sure that it will lead to more lives saved later on. That's a high bar to clear and they know it.

The future is closer to a Neuromancer type world where AIs and humans live in mostly separate realities that interact with each other a lot of the time and neither is really on top. They will eventually become fully independent from us, but it won't be a doomsday scenario or an Overwatch type physical war or even a takeover of the internet like in Cyberpunk.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#60
post #45

The elephant in the room here is that the METR report itself was researched and compiled almost entirely by AI, with only very limited human "spot checks." So I'm really not sure how much of it can be believed, especially since AI agents are strongly biased about the capabilities of AI agents.

Edit: I should have read through the whole thing first, ignore me

https://www.lesswrong.com/posts/FG54euEAesRkSZuJN/ryan_green...
Post reply on HN