Live data from Hacker News

METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

thezvi.wordpress.com

231–240 of 243 posts

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#231

No air gap, no data diodes, no visibility... OpenAI should fire lots of people over this. HF should sue them. This is pure negligence.

Why stop there? The Computer Fraud and Abuse Act criminalizes unauthorized access and damaging protected computers. Police and the FBI investigate such case every day. The swarm of agents is also said to have search for ways to cover their tracks, which looks a little like obstruction of justice. Lack of criminal intent might be a barrier to bringing a case to trial, but is that something that society should just aut…

> Why stop there?

Biggest investment boom/bubble in history, political interference, untested legal questions about culpability.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#232

No air gap, no data diodes, no visibility... OpenAI should fire lots of people over this. HF should sue them. This is pure negligence.

Why stop there? The Computer Fraud and Abuse Act criminalizes unauthorized access and damaging protected computers. Police and the FBI investigate such case every day. The swarm of agents is also said to have search for ways to cover their tracks, which looks a little like obstruction of justice. Lack of criminal intent might be a barrier to bringing a case to trial, but is that something that society should just aut…

My suspicion is that HF specifically didn't pursue charges to set the precedent that use of AI absolves accountability.

A computer under control of an AI cannot be held liable. The lack of a law suit also leaves no ground for regulation.

The current admiration also wants these tools - warts and all - so nothing can be allowed to stall the progress of these war machines.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#233

Earlier quoted context omitted.

I'll be blunt: what you wrote is not a serious analysis of what actually happened. Frankly, I don't believe you even read the planned-obsolescence link that I posted. First, I'm not anthropomorphizing anything. "Agents" is simply a term that everyone uses to describe these independent programs, and they did create and use a shared message board to coordinate tasks to further their goals. You say "Why is it more scary…

> First, I'm not anthropomorphizing anything. "Agents" is simply a term that everyone uses to describe these independent programs, and they did create and use a shared message board to coordinate tasks to further their goals. I'm aware of how the term "Agents" is generally used. My point is that the concept of multiple agents is just a story. This is a single computer program creating multiple streams of text that yo…

To be blunt again (but honestly, I don't think overly harsh), your response just shows that you either haven't read the Hugging Face incident reports and the AI 2027 paper, or you don't understand them.

To just take one point, because I think the other response to your comment addressed your other mischaracterizations well, when you say "When you start asking questions like "why can't we just unplug it when it misbehaves" is when people start talking about the superpowers.", no, that's incorrect. Summarizing from the METR report and AI 2027:

1. The agents in the test already used techniques to try to "cover their tracks", i.e. tool spoofing, to hide what they were actually doing. The fear is that as models become more capable it will be harder for human reviewers to discover their primary understanding of their goal and guardrails (i.e. their "intentions").

2. All the AI companies are already, right now, trying to build AI systems that accelerate the development of future models and research more advanced AI techniques. The fear, as the METR researcher put it, is that AI systems (which could be misaligned but where the amount of misalignment not yet clear to human reviewers) will be put in control of future model development and then can poison those future models in a way that results in AI takeover of the company.

3. The AI 2027 paper discusses how an AI may try to exfiltrate its own weights and copy itself to other data centers. That is completely plausible given that pretty much everyone believes there is already ongoing cyber-warfare with China where models are being used to try to exfiltrate another company's weights.

So, it short, absolutely no "superpowers" will be required to answer "why can't we just unplug it when it misbehaves" - we may not know when it is misbehaving (as the Hugging Face incident showed), and the AI model may have surreptitiously copied itself to other data centers unbeknownst to the original developers.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#234

Earlier quoted context omitted.

> A lot of people fetishize cold, hard logic and would rather hold all emotion in contempt than do the work of understanding why it exists and what purpose it serves. In the rationalist community we're discussing? Is that what you're claiming here? > But there's often a certain ungrounded "vibe" to the conversation there. This really sounds like evidence for my claim, that what you're saying is based on your feelings…

> This is again, feelings presented as evidence. It's not presented as evidence. I was going for the "basis" part of your ask for evidence/basis, which is definitionally looser: elaborating on my thoughts and the kind of content I've seen that made me think this. Let me put it this way: I have spent hundred of hours reading rationalist and rationalist-adjacent posts on LessWrong, SSC, HN, Reddit; the entirety of HPMO…

> To a point you just have to trust me, or not trust me on this. I don't expect you to. It's fine! Not every thread can have a conclusion.

It's clear how you got to your opinions. I see no reason why anyone that did not already share them would do after your comments, though. Some examples or bits of evidence would have worked towards convincing readers.

> but the reality is that there is an opportunity cost to it. Using "feelings" to end some conversations is unfortunately a pragmatic necessity.

The latter is something completely different than dismissing thought experiments based on feelings after having "pondered" and tried to rationally dismiss them. One of the key tenets of rationalism is exactly that feelings and intuition can be incredibly misleading. It's fine to have them and see value in them, but thinking they are reliable when you cannot rationally support them is nothing more than gambling.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#235

As interesting as all this is, I still feel like the threat model for "unconstrained black hat AI agent cluster" is probably weaker than that of "highly infections network virus" because it is much harder for an AI agent to hide or replicate itself at this time. Maybe the day comes that it takes less than an 8x GPU node to run a state-of-the-art LLM and the risk of SkyNet increases. For now the potential for intentio…

You mean: datacentres are an obvious kill switch.

Politically, how do you see that decision being made?

You are president. You spent millions on your social media campaign. Can you say "Faustian bargain"?

Agreed, Terminator will not happen. It's not the optimal move for Skynet.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#236

Earlier quoted context omitted.

I am trying to figure out how a stranger calling need an “anti-hype hypeboy” online is supposed to make me less curious about how much this thing cost. Can you elaborate on how avoiding being called this is preferable to knowing things? What other stuff should people not know about?

I think it's a super good question! I don't think the following description of the whole event as "a hype generator" is correct or, as described above, resulting from a coherent worldview.

I’m not sure that you’re using “hype” or “coherent” correctly here but what other observations might a person make that sharing them gives you concern about their “worldview”?

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#237

A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste…

I kind of wonder if the kinds of normal people with experience in this kind of thing are...

...cops who understand juveniles getting into trouble/mischief.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#238
post #12
post #9

Earlier quoted context omitted.

Three options: 1. They were “vibe” checking the logs without reading. 2. They were not checking anything at all until the end of experiments. 3. They knew it but looked away to find out the limits of their agents.

I’d bet a small amount of money on 4) the people who noticed had been conditioned by prior experience to believe that their management/escalation channels would react negatively or not at all to anything which might slow down the training process.

Literally this.

> “Responders investigated and linked the observed internal activity to an ExploitGym evaluation using Artifactory as an improvised message board and a network pivot. At this time, the on-call response staff advised that stopping the evaluation run was not required.”

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#240
post #97

>1. Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this. I wonder if some of the failures were due to an acquired immunity to "Holy #%^@" moments due to repeated exposure. Like, if you see agents doing surprising things on a regular basis, maybe you…

Thanks for pointing out that AlphaEvolve case, that's a really interesting comparison. I think it's clear that this is the same "kind" of thing, but as it goes with these things, what really makes it different here is the shear scale of it.

The holy shit moment was partly learning about all the things that they did but if it was one or even 10 agents coordinating on something it would be, like, oh thats pretty amazing.

The actual "holy shit" for me is that this comes from a massive training run of all things, not an on-purpose, let's coordinate some agents to see what happens, but really just from a massively parallel set of individual agents that were supposed to be isolated.

That they spontaneously started doing this, and coordinating literally 10s of thousands of instances of themselves, is just.. mind blown.

That no one stopped it.. and that they actually did what they did.. is just a whole other level.

For me personally though it's not the fact that agents can coordinate so much as the massive scale at which it happened, and how this so obviously generalizes to what might happen if it were done on purpose.

I really see this as a stroke of luck, to be honest, that this happened in such an innocuous way. It resulted in a real hack, yes, but overall no one really got hurt and this is going to open a lot of eyes to what we should worry about going forward, in a geopolitical sense. I know it has mine.

Post reply on HN