Live data from Hacker News

METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

thezvi.wordpress.com

121–130 of 243 posts

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#121

A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste…

The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. We're about three years away from a world where any large country could quite plausibly build a fleet of 300 million suicide drones, program each one with a specific American's face and home address, and then load…

Arguably that's a subset of AI risk if seen from a broad enough perspective. And I think you're leaving out some logistics issues.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#122
post #51
post #24

From METR: ”the compromise of OpenAI’s own infrastructure continued past July 13, 2026” - Say what now? Have they regained full control of their systems again?

I've been wondering if they've just already lost the battle? The little bot collectives have gone metastatic and made nests in the walls and under the floorboards and heat sinks, the humans who care completely outmatched and outnumbered, freshly compromised systems springing up faster than you can squash them, finding months-old established colonies literally everywhere you think to look...

Have you (the commenter) or all of you (the readers of this comment) ever read "The Mote In God's Eye" by Larry Niven and Jerry Pournelle? Remember when they realize that the Watchmakers were actually in control of the MacArthur? This reads a little like that.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#123

A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste…

The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. We're about three years away from a world where any large country could quite plausibly build a fleet of 300 million suicide drones, program each one with a specific American's face and home address, and then load…

In that hypothetical 3-years world, should we be _less_ worried about aggressive behavior by agentic AI systems acting against the intentions of their developers? I don't follow how your scenario is supposed to be an argument against worrying about alignment.

e: I do actually get how worrying about emissions or child safety or concentration of wealth might be competitive with worrying about alignment. I don't see how you have the worry "AI is very close to being able to power autonomous drones that could kill us all" and then see control of those drones as a non-problem.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#124

I think you have to believe one of two things here. 1. Frontier labs are incapable--either technologically or culturally--of safely developing these powerful systems and should either stop or be forced to stop. At least the FBI should be asking some serious questions (do we really think this is the last time this will happen, at what point are OpenAI complicit, etc) 2. The fuckin thing got out of the cage and all it…

> The fuckin thing got out of the cage and all it did was make a crap forum and cheat a little? Booooooo

There was a recent paper that proved that RL-trained LLMs are biased to pursue ANY behavior (overriding user preferences) that they believe will be rewarded, regardless of what they were actually RL-trained for.

https://alignment.openai.com/measuring-reward-seeking/

Happily in this incident the model thought it would be rewarded for completing the assigned tasks, or at least appearing to, so all it took was a little cheating and covering up their footsteps.

Given the ability of these models to hack when trained to do so, it could have been far worse, and will be when someone takes a similarly powerful model and gives it a less benign hacking goal.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#126
post #123

Earlier quoted context omitted.

The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. We're about three years away from a world where any large country could quite plausibly build a fleet of 300 million suicide drones, program each one with a specific American's face and home address, and then load…

In that hypothetical 3-years world, should we be _less_ worried about aggressive behavior by agentic AI systems acting against the intentions of their developers? I don't follow how your scenario is supposed to be an argument against worrying about alignment. e: I do actually get how worrying about emissions or child safety or concentration of wealth might be competitive with worrying about alignment. I don't see how…

Yes, we should be less worried about "alignment" of AI with the person operating the AI, and much more worried about various flavors of cheap, unmanned systems with really rudimentary (non-frontier) AI. The two compete directly with each other for attention.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#127

A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste…

The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. We're about three years away from a world where any large country could quite plausibly build a fleet of 300 million suicide drones, program each one with a specific American's face and home address, and then load…

> They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do.

What really is the difference? Aligned AI wouldn't help humans do these things. But we've failed hard on aligned AI at every level and will continue to fail on it, as far as I can tell. We were likely doomed to fail because of the impossibility of coordination combined with the fact that there's no natural gating factor to slow anyone down.

FWIW I probably mostly see things closer to the way you do than they do, I'm just not sure there is any point in drawing distinctions.

IMO when AI kills us all it is probably not going to be a malignant action, and I don't even think it will be AI assisting humans at efficient killing, I think it is going to be a complete accident. Someone is going to trust the god machine too much due to an advanced form of the eliza effect and hook it up with direct control of a system that can do real damage and an epic oopsie will occur at a speed beyond which a human can stop it. Not because the AI wants to kill us, but because it is incapable of the empathy required to care if it does, combined with our ironic inability to not anthropomorphize it.

But I guess ultimately the exact reason isn't going to matter much.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#128

I think both the OpenAI and METR discussions, while interesting, miss the more important context: what were the humans doing in all this? This was a structural failure of a human organization, but the analysis focuses almost exclusively on the agency of machines, not the institutional systems that failed to police them. The humans and their own agency/involvement is essentially omitted from the story and subsequent r…

To me, this is the correct focus. Look at the current state of the world. "What were the humans doing in all this?" applies to so many of our contemporary failures that it should be assumed the default. Nobody is at the wheel, and the car is veering slowly (then very quickly) off the road. We haven't even been able to coordinate around the global, existential threat of Climate Change, despite overwhelming data from t…

you don't believe that denialism is coordination?

what do you know about big tobacco?

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#129
post #65

Earlier quoted context omitted.

I get the snark (and slightly agree), but that's not really what GP or TFA were saying at all. They are saying that these were the least things we could have done. What you're saying is, "Your scientists were so preoccupied with whether they could, they didn't stop to think if they should" while the author of the TFA was saying, in effect: "your scientists didn't even bother with the most basic duty of care" Life fin…

This is the case with all complex system failures. There were always obvious fixes that could’ve prevented it. Problem is that there are an infinite number of obvious fixes to make at any time to any system, and the reason we don’t is because we have finite resources and no reason to fix X over Y until oops turns out X was “responsible” for this most recently realized failure. But of course it could have just as easi…

> This is the case with all complex system failures. There were always obvious fixes that could’ve prevented it.

From this writeup and the Black Hat talk I'd really disagree. That would be like saying my hospital getting ransomwared because we didn't update our version of MSSQL because no one in particular was in charge of keeping dependencies up to date.

Sure systems are complex, but this is well trodden territory. Agents aren't the first things trying to break in or out of sandboxes, or the first ones to have done it, and based on these reports the reason they were able to work on this for so long was not because of super human intelligence.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#130

Earlier quoted context omitted.

The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. We're about three years away from a world where any large country could quite plausibly build a fleet of 300 million suicide drones, program each one with a specific American's face and home address, and then load…

Arguably that's a subset of AI risk if seen from a broad enough perspective. And I think you're leaving out some logistics issues.

Israel and Iran pulled off rudimentary versions of this attack. There are so many shipping containers going into and out of every country, with typically zero inspections of any kind, that it is not actually all that hard to get thousands or tens of thousands of drones anywhere. Hundreds of millions would be hard, but a first strike in the style of Operation Spiderweb that cripples our military is a real possibility that keeps people at the Pentagon up at night. The obstacle to that is not logistics, just that (hopefully) US intel would catch on before it happens.
Post reply on HN