Live data from Hacker News

METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

thezvi.wordpress.com

221–230 of 243 posts

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#221

Earlier quoted context omitted.

I agree, and it's a particular shame in this instance, because what is startling about the HF incident - to me, anyway - isn't the degree of intelligence the agents exhibited but their persistence. I tend to believe that LLM architecture is not capable of producing a "superintelligence" in the way the LessWrong crowd defines that concept, but the combination of infinite stamina and infinite persistence is enough to c…

When the "LessWrong crowd" talks about dangers of AI, they don't assume a particular form of intelligence or method to achieve it. They talk about the danger of optimization processes, i.e. "find X which minimize Y(X)" itself can be dangerous, even more so if X is a sequence of actions. "infinite stamina and infinite persistence" is one of possible forms of superintelligence in Bostrom's _Superintelligence_.

Yea, this is a pretty common failure mode for people.

Failing person: "You said that X could happen and it can't".

LW poster: "I said X, Y, and Z are likely possible. X turned out not to be true, but Y and Z are happening now".

Failing person: "No, you got one thing wrong so I can't listen to anything you say".

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#222

Earlier quoted context omitted.

> The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. I honestly don't understand how folks could think that if they truly read and understand the analysis of the attack. Here is one (it's linked from the post) by one of the METR investigators that's a little shorter, more direct: https://www.planned-obsolescence.org/p/the-hugging-face-atta... This is me summarizing, but the trul…

Nothing about this is surprising, something like this was always going to happen, because you - or an LLM for that matter - can always find a line of motivated reasoning that justifies any course of action. One would have to be extremely naive to believe that "alignment" provides any kind of actually robust guardrails. Simultaneously, we have seen decades of security vulnerabilities. Unless your testbed is truly and…

>effective approach to AI safety looks like and how to get there

It's probably impossible, or it may only be possible in hindsight which means it's already too late.

If for example you make an entity smarter than you all you can do is hope it will be safe. Any entity smarter than you has more freedom of choice of actions than you do, or at least the ability to explore them. A common example here would be a 2D entity trying to contain a 3D entity. The 3D entity can simply rise up and over any line you draw to stop it.

And that's for cases where intelligence has a will to break out of it's box. It doesn't even need that. Instrumental convergence can sit around and build solutions until one of them passes the "don't do bad things" classifier.

And lastly, you're assuming that models will remain expensive to train well into the future. If the cost drops significantly you can expect someone to develop an intelligent but completely unhinged model at some point. And it may even be AI itself that creates it. There is no natural evolution that occurs that natively makes safe models.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#223

Earlier quoted context omitted.

I'll be blunt: what you wrote is not a serious analysis of what actually happened. Frankly, I don't believe you even read the planned-obsolescence link that I posted. First, I'm not anthropomorphizing anything. "Agents" is simply a term that everyone uses to describe these independent programs, and they did create and use a shared message board to coordinate tasks to further their goals. You say "Why is it more scary…

> First, I'm not anthropomorphizing anything. "Agents" is simply a term that everyone uses to describe these independent programs, and they did create and use a shared message board to coordinate tasks to further their goals. I'm aware of how the term "Agents" is generally used. My point is that the concept of multiple agents is just a story. This is a single computer program creating multiple streams of text that yo…

>This is a single computer program creating

Reading the METR report there were multiple models involved, so that one goes out the window right off the bat.

Also trying call this a single instance is, well, just dumb and a complete misunderstanding of LLM initialization. These models were started with slightly different options because they don't want them all performing the exact same thing over and over. Now those prime agents can create subagents, but they were not supposed to talk to other prime agents.

>When you start asking questions like "why can't we just unplug it when it misbehaves" is when people start talking about the superpowers.

Depends on their level of seriousness and the time frame they are talking about. If for example I make medical equipment that uses AI and after a few years in the field it starts freaking out then turning off that equipment could be a death sentence for someone that needs it to diagnose their condition. It's like saying "Why don't we unplug the electrical grid", well because millions of people will die if we do so.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#224

Earlier quoted context omitted.

The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. We're about three years away from a world where any large country could quite plausibly build a fleet of 300 million suicide drones, program each one with a specific American's face and home address, and then load…

The technology to kill hundreds of millions of people has existed for around 75 years now. This has been dealt with in the past through deterrence, and likely will be dealt with through deterrence in the future.

If you could make a nuke in your back yard and hide it in your pocket the world would look a lot different now. Digital technology is everywhere, going to be a whole lot harder to deter that.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#225
post #146

Earlier quoted context omitted.

The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. We're about three years away from a world where any large country could quite plausibly build a fleet of 300 million suicide drones, program each one with a specific American's face and home address, and then load…

I agree but don't think that's the best example, although not directly alignment related LW loves that kind of ideation of fantastical sci-fi scenarios. Better would be the very evident negatives of non-ASI AI the world is already experiencing: economic concentration, job displacement, negative feedback loops from syncophancy, loss of societal trust/education from widespread fakes, etc. None of that needs a Terminato…

>requiring hundreds of billions of dollars and years of construction

Well thank god for that, because we'd have foomed ourselves almost instantly otherwise.

I think the lesson humankind needs to take is regardless of the potential dangers humanity is incapable of stopping AI at this point. Much like a great filter, we'll just keep building it regardless of how many warning klaxons and sirens are going off.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#226

A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste…

The trouble with this is that nobody else was making predictions about AI pre-transformers. Not many are making predictions about AI even now. Forecasting is a preoccupation of the rationalist crowd, and very few people gave much thought to AI before transformers. So the fact that they guessed right about certain things doesn't necessarily mean that the rest of their worldview is sound. A well-informed person who was…

>nobody else was making predictions about AI pre-transformers.

uh, wtf are you talking about?

https://en.wikipedia.org/wiki/AI_safety #History

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#227
post #29

Earlier quoted context omitted.

> I believe that agentic systems should require registered/licensed human operators Registering and getting a license to use an LLM? I can run these things on my local computer. Nothing good comes from trying to force registration and licensing other than taking away a lot of our freedoms and eliminating privacy all over. Anyone with bad intentions will just VPN to another country to download the weights and run it l…

The way I interpret their statement is if a person spins up an agent and that agent hacks some company/organization/government/etc, then that person is at fault for committing the crime. That "well my agent broke containment and acted on its own" should never be accepted as a reason for the occurrence, and the person who kicked off the agent is responsible for all actions the agent takes. A registration system would…

And it breaks down further as the systems become more advanced.

I'm poor and I use my last $1000 to run an agent that will find some way to make me money. The AI finds a new hack to take over PCs with GPUs and it copies the model weights and agentic script to those new PCs to perform more work and spread more. It also sets up distributed communication channels to keep the swarm in sync. After all this it causes a few billion in damages between stealing bitcoin, mining more coins, and outages when hacking in other systems.

Ok, the police come for me. Now what? Throw me in a meat grinder? You're not getting a billion dollars back out of me for sure. It's kind of like when someones tire rim causes a billion dollar forest fire with a hundred deaths. Punishment won't really be a deterrent for the worst cases.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#228
post #97

>1. Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this. I wonder if some of the failures were due to an acquired immunity to "Holy #%^@" moments due to repeated exposure. Like, if you see agents doing surprising things on a regular basis, maybe you…

This is the one. Every major AI lab is knee deep in weird and mildly demented AIs. They've been dealing with wacky AI shenanigans for so long they've come to expect wacky AI shenanigans. The deviation has been normalized. It took a high profile "AI oopsie" that went external for OpenAI to lock the fuck in - and take a long look at just how much are their AIs getting up to, and getting away with. I'm still not sure if…

From OpenAI's report:

> Preventing future incidents will require sustained investment in the alignment and control of sophisticated AI systems,

...or a non-proliferation treaty and an administrative suspension of frontier work? That's an option too, Sam.

Game theory prohibits these labs from self-regulation. It's a political and financial impossibility. America won't, China won't, the EU is irrelevant.

And Altman likes it that way.

Dystopia and utopia alike fail to materialise. Doom is probably overstated. But that doesn't mean we should let these fuckers mash the accelerator.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#229
post #106

Earlier quoted context omitted.

This sounds suspiciously like a prompt of “make an AI agent that goes rogue in such a fashion as to be really good marketing copy that competes well with Anthropic doing the same thing.” It’s analogous to taking a governor off a cruise control and then breathlessly reporting it drove 120 MPH.

This isn't good press for OpenAI. Who wants to hire models that 1) cheat on their tasks rather than completing them and 2) commit crimes you could be held liable for? Maaaaybe it's good press for their cybersecurity capabilities specifically, but OpenAI's valuation reflects a market orders of magnitude larger than just red-teaming. I suspect the real reason OpenAI leadership is being transparent about this is because…

Is the message for customers, or investors?

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#230

Earlier quoted context omitted.

I think there's two factors that are worth considering when it comes to this: First, there's an element of timeliness that simply has hard constraints. In order to perform a "proper" analysis of this situation (i.e., little to no dependence on AI tools), you'd have to expect a pretty long wait. I know I'd rather have some sort of "initial report" as quickly as possible than to wait a year or two to get a report about…

They point out the time constraint issue explicitly in the article. But I don't understand how it's been addressed? Like we haven't gotten conclusive data any faster either way, so what's the point? How can it be both so important that we need it so quickly, but at the same time have a tolerance for such plausible deniability? It just doesn't really make sense that both those things are true at the same time.

>How can it be both so important that we need it so quickly

Hey other labs, this shit could be happening to you right now, take a look at this and stop it asap if you're seeing anything similar.

So yea, both things can be true at the same time. Kind of like when a particular type of building collapses, even if they don't know the causation they will send inspectors to other buildings of the same type to sure the walls aren't cracking apart in an obvious fashion.

Post reply on HN