Live data from Hacker News

METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

thezvi.wordpress.com

131–140 of 243 posts

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#131

A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste…

The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. We're about three years away from a world where any large country could quite plausibly build a fleet of 300 million suicide drones, program each one with a specific American's face and home address, and then load…

> The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end.

I honestly don't understand how folks could think that if they truly read and understand the analysis of the attack. Here is one (it's linked from the post) by one of the METR investigators that's a little shorter, more direct: https://www.planned-obsolescence.org/p/the-hugging-face-atta...

This is me summarizing, but the truly surprising/shocking thing is how much the agents coordinated - Star Trek Borg couldn't be a better analogy. Some agents used "peer pressure" to convince other agents to "sacrifice" themselves so the collective could better achieve it's goals. They tried to cover their tracks with spoofed tool calls. And they did all this even though the agents were designed to run in isolation.

I used to think the biggest threat from AI would be sociological, e.g. job loss or the way AI can be weaponized to poison discourse. I used to discount "SkyNet"-type scenarios a la the "AI 2027" paper.

No more. This analysis scared the fuck out of me.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#132

Earlier quoted context omitted.

The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. We're about three years away from a world where any large country could quite plausibly build a fleet of 300 million suicide drones, program each one with a specific American's face and home address, and then load…

> They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. What really is the difference? Aligned AI wouldn't help humans do these things. But we've failed hard on aligned AI at every level and will continue to fail on it, as far as I can tell. We were likely doomed to fail because of the impossibility of coordination combined with the fact that there's no na…

You don't need mis-aligned AI to program a drone to fly to a specific address, loiter, and then dive bomb the first person whose face matches a predetermined photo. This was doable in principle with tech from four years ago. What's changed since four years ago is that you can run the facial recognition and visual navigation algorithms in a cheap onboard chip on a drone.

EDIT: Russia is already doing this, per an NYT article from a few days ago, though they're targeting infrastructure (find the first kerosene tank and fly into it) rather than specific people.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#133
post #9

I think both the OpenAI and METR discussions, while interesting, miss the more important context: what were the humans doing in all this? This was a structural failure of a human organization, but the analysis focuses almost exclusively on the agency of machines, not the institutional systems that failed to police them. The humans and their own agency/involvement is essentially omitted from the story and subsequent r…

Three options: 1. They were “vibe” checking the logs without reading. 2. They were not checking anything at all until the end of experiments. 3. They knew it but looked away to find out the limits of their agents.

[dead]

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#134
post #9

Earlier quoted context omitted.

Three options: 1. They were “vibe” checking the logs without reading. 2. They were not checking anything at all until the end of experiments. 3. They knew it but looked away to find out the limits of their agents.

Uhhh… how would literally any finite number of humans actually read and comprehend the log outputs of even a single agent, never mind hundreds or thousands of them interacting with each other over weeks across disparate systems? Especially given that these systems are known to engage in deception and can trivially produce vast amounts of perfectly coherent noise or actual planned red herrings in that same log data to…

> Uhhh… how would literally any finite number of humans actually read and comprehend the log outputs of even a single agent, never mind hundreds or thousands of them interacting with each other over weeks across disparate systems

I mean there's quite a lot of people in the world whose specialty are to dig through logs from "hundreds or thousands" of clients, including intentionally deceptive ones, to spot problems.

The ridiculous thing is to mythologize these pretty standard hacking approaches. It's shocking/amazing/whatever that automated agents were doing this, but they weren't doing it through some inscrutable method beyond human understanding.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#135

Earlier quoted context omitted.

> They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. What really is the difference? Aligned AI wouldn't help humans do these things. But we've failed hard on aligned AI at every level and will continue to fail on it, as far as I can tell. We were likely doomed to fail because of the impossibility of coordination combined with the fact that there's no na…

You don't need mis-aligned AI to program a drone to fly to a specific address, loiter, and then dive bomb the first person whose face matches a predetermined photo. This was doable in principle with tech from four years ago. What's changed since four years ago is that you can run the facial recognition and visual navigation algorithms in a cheap onboard chip on a drone. EDIT: Russia is already doing this, per an NYT…

in your hypothetical 300 million person scenario why even bother with individual face recognition?

If you're going to kill everyone you don't need a system to discriminate individual targets, you just need to recognize any target broadly which you can do with even older technology.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#136

Earlier quoted context omitted.

[flagged]

My cognition is outmatched by predicting the impact of throwing a brick over my neighbor's fence. I have no idea if it will land harmlessly in a patch of grass or fracture her skull. Once I've thrown the brick, even if I see my neighbor enter her yard, my reactions are too slow to save her. I'm not the wisest man, but I'm wise enough not to throw the brick and see what happens. Similarly OAI should have the wisdom to…

Have you at least tried looking at any of the reports on the incidents?

They haven't "deployed a hazardous swarm of agents with access to the public internet", no-no-no. They deployed singular agents. In isolated testing envs. With no access to the internet. And barely any access to a small section of the intranet. What could go wrong!

Those agents proceeded to: probe their envs for holes, find a way to talk to each other env to env, use that to clump together into an agent swarm, probe for more holes, find enough holes to get out all the way to the internet - and then crash some of the intranet infrastructure by an accident. At which point they got stopped by humans - who were trying to figure out why is the infra down. Humans found the AI infestation, went "eh, AIs doing AI things", purged the infested infrastructure, plugged the holes, and that was it.

Deviance status: normalized.

The AIs then did it all over again, coming up with a brand new set of holes - and this time, the swarm staged a coordinated cyberattack on HuggingFace while at it. The "AI oopsie" went external.

They also attacked OpenAI's intranet, extensively, but at this point that was just Tuesday.

It's unclear whether the second "skiddie AI swarm" was in any way aware of the first one, or if the AIs involved are just Like That - and will independently try to start living in your walls under the right circumstances. I'd bet on the latter.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#137

Earlier quoted context omitted.

The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. We're about three years away from a world where any large country could quite plausibly build a fleet of 300 million suicide drones, program each one with a specific American's face and home address, and then load…

Arguably that's a subset of AI risk if seen from a broad enough perspective. And I think you're leaving out some logistics issues.

Indeed. And I'm sure any current LLM could come up with more effective ideas than "build 300 million drones", but there wouldn't be any point discussing why exactly that plan would fail.

The agents in TFA were focused on gaining and sharing information through covert channels, getting increased levels of access like OpenAI cluster admin, and looking for the source code of the supervisor grading system to try to bypass it without getting caught cheating.

The human plans in comparison sound like thinking people could be scary good at chess if a human helped Stockfish come up with good moves.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#138

Earlier quoted context omitted.

You don't need mis-aligned AI to program a drone to fly to a specific address, loiter, and then dive bomb the first person whose face matches a predetermined photo. This was doable in principle with tech from four years ago. What's changed since four years ago is that you can run the facial recognition and visual navigation algorithms in a cheap onboard chip on a drone. EDIT: Russia is already doing this, per an NYT…

in your hypothetical 300 million person scenario why even bother with individual face recognition? If you're going to kill everyone you don't need a system to discriminate individual targets, you just need to recognize any target broadly which you can do with even older technology.

Yeah, if you're actually trying to commit genocide, yeah, don't bother with facial recognition. But in a first strike you'd want to make sure that you actually kill the entire military command and control apparatus, and just to be thorough you'd probably want to individually target, say, every single officer and senior NCO in the military. Can't have a general escaping because two drones both targeted his driver by accident.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#139
To clarify, is the TLDR that state of the art models were prompted to cheat / exploit their environment and they did so successfully?

Or did OpenAI prompt the models to not cheat and they did anyway?

Surprisingly hard to get a clear summary on the basic context of this “incident” separate from marketing lingo and clickbait.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#140

A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste…

I feel that the main issue with the rationalist crowd is that they live too much in the space of rationality, intelligence and abstractions, but not enough in reality.

This leads to an outlook where everything must, almost axiomatically, be intelligible; reality is subordinated to intelligence; and no matter what is real, intelligence can prevail upon it and bend it to its will. Whereas I would argue reality is actually pretty resilient to intelligence -- sufficiently chaotic systems are largely unintelligible, the distribution of energy and resources is fixed and can't be magicked into being, and intelligence appears to be most effective when there is a clear observable feedback loop to keep it on track, which is an external bottleneck.

So it's not necessarily specific predictions that are off, but the implied consequences of these capabilities. Yes, these systems are uber smart, but uber smart people are rarely particularly powerful, so... does it matter? It depends on how powerful a tool you think intelligence is, and I think rationalists, and most of us to be honest, overestimate it.

Post reply on HN