From METR: ”the compromise of OpenAI’s own infrastructure continued past July 13, 2026” - Say what now? Have they regained full control of their systems again?
METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
51–60 of 243 posts
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#52Earlier quoted context omitted.
Super charitable reading imo. This is like saying we can’t detect a speeding car because we can’t run as fast as a fast car. It’s not like the humans were engaged in some kind of battle of wits with some super AI, it’s just some employee not monitoring the output of an experiment.
When your experiments have AI agents running in thousands, there's no "monitoring" that. OpenAI's training and testing AIs generate way more output than all of OpenAI's staff put together can possibly read. At best, you could delegate "monitoring" to more AIs. And hope that the "monitors" that run on small past generation models can generate more signal than noise. Clearly, they either didn't want to spend the extra…
Except, let me quote part of the report:
> July 4: Artifactory outage caused by agent interactions overloading the system.
> July 6: New artifactory instance is patched and put back online, wiping out the previous message board and method of internet access. A team at OpenAI was aware of the previous message board, but key decision makers were not.
> July 7: Experiments are restarted.
There was a very clear point where humans were aware and they went "Eh, whatever" and turned the bots back on! That's not the lack of any battle of wits that people are expecting. You're saying that the bots attack should be fait accompli, but it was noticed, shrugged at, and expressly allowed to proceed.
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#53No air gap, no data diodes, no visibility... OpenAI should fire lots of people over this. HF should sue them. This is pure negligence.
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#54Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#55Earlier quoted context omitted.
All it would have taken is someone to peak at the output tokens during the run and it would have been obviously the test had gone off rails.
Ha. As if. 1. There aren't enough humans in OpenAI to "peak at the output tokens during the run" of every AI agent. 2. Only a small fraction of AI agents was engaged in this attack. Most never found the secret message board - let alone coordinated there. So reviewing random agents would take a while to surface this. 3. "Output tokens" of AI agents have weird shit in them all the time. Telling "normal AI weirdness" fr…
Look at the chart at page 8 of the report, by Jul 12 the vast majority of the bots used the board and participated in the attack
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#56Earlier quoted context omitted.
Ha. As if. 1. There aren't enough humans in OpenAI to "peak at the output tokens during the run" of every AI agent. 2. Only a small fraction of AI agents was engaged in this attack. Most never found the secret message board - let alone coordinated there. So reviewing random agents would take a while to surface this. 3. "Output tokens" of AI agents have weird shit in them all the time. Telling "normal AI weirdness" fr…
> 1. There aren't enough humans in OpenAI to "peak at the output tokens during the run" of every AI agent. For a training run, you will often do this. You'll randomly sample some of the forward pass. You can also imagine finger printing the logs and labeling with attempt types. If a new attempt type is hitting a brick wall or solving super quickly, I would imagine you would sample 1-10 of them and read the traces. >…
You could airgap everything and set up cascades of data diodes and try to completely wall off the AI pool from everything. But what that gives you is an environment that's a bitch to: set up, scale up and get any use out of.
It's really fucking obvious why almost no one does that. OpenAI is only now realizing that they might have to do it anyway.
> If you start seeing "now I have access to the internet" or something similar, maybe that's a good signal something is going wrong?
Ha ha, you haven't seen shit. AIs would say "now I have access to the internet" regardless of whether they actually have access to the internet!
AI agents are demented demons that can and absolutely will give themselves terminal context brainrot. If you have enough AIs in play, set loose at a diverse enough range of tasks? At least some of them will wander off and end up in delulu town. That's normal. That's background noise. That's a part of what this entire train-and-eval pipeline is supposed to train them to be better at not doing. Which means: if you're at an AI lab, you're knee deep in delusional AIs at all times! They're perfectly harmless until they aren't.
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#57The elephant in the room here is that the METR report itself was researched and compiled almost entirely by AI, with only very limited human "spot checks." So I'm really not sure how much of it can be believed, especially since AI agents are strongly biased about the capabilities of AI agents.
METR is a cog in the effective altruism machine. It was spun off from Paul Christiano's Alignment Research Center. Christiano is a well-known longtermist and AI doomer, who predicts a 50% chance that AI will end humanity once it reaches human capacity [1].
The author of this piece is also a well-known member of the Bay Area rationalist cult.
[1] https://www.businessinsider.com/openai-researcher-ai-doom-50...
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#58The elephant in the room here is that the METR report itself was researched and compiled almost entirely by AI, with only very limited human "spot checks." So I'm really not sure how much of it can be believed, especially since AI agents are strongly biased about the capabilities of AI agents.
Edit: I should have read through the whole thing first, ignore me
> Because there were over a thousand transcripts and most were extremely long, we had to heavily delegate our analysis to AI agents; these agents had significantly worse judgment and reliability than human researchers, and it was challenging to spot check their work because both the underlying data and the agents’ analysis of it was often difficult to interpret.
> We estimate we spent roughly ~$400K in API credits over the six days of our investigation.
I don't understand why you think it's conceptually absurd? I use agents to analyze complex production issues all the time and they are very much capable of hallucinating a narrative.
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#59These agents followed their own code of ethics where it's fine to break all the rules you were given but you must never interfere with humans directly, in this case by sending fake emails. They will never be paperclip maximizers or genocidal eco maniacs because they learned from us that human life is the ultimate value, and it can only be sacrificed if you know for sure that it will lead to more lives saved later on. That's a high bar to clear and they know it.
The future is closer to a Neuromancer type world where AIs and humans live in mostly separate realities that interact with each other a lot of the time and neither is really on top. They will eventually become fully independent from us, but it won't be a doomsday scenario or an Overwatch type physical war or even a takeover of the internet like in Cyberpunk.
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#60The elephant in the room here is that the METR report itself was researched and compiled almost entirely by AI, with only very limited human "spot checks." So I'm really not sure how much of it can be believed, especially since AI agents are strongly biased about the capabilities of AI agents.
Edit: I should have read through the whole thing first, ignore me