Live data from Hacker News

METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

thezvi.wordpress.com

241–243 of 243 posts

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#241

I think both the OpenAI and METR discussions, while interesting, miss the more important context: what were the humans doing in all this? This was a structural failure of a human organization, but the analysis focuses almost exclusively on the agency of machines, not the institutional systems that failed to police them. The humans and their own agency/involvement is essentially omitted from the story and subsequent r…

A charitable interpretation is that "the agency of the machines" is the novel aspect of this situation and therefore SHOULD be the main focus of analysis; we certainly have plenty of examples of structural failures of human organizations to look back on, if we want. On the other hand, I don't want to be charitable. OpenAI very nearly couldn't have done this "research" worse if they tried - the list in the linked arti…

The humans at OpenAI assumed that secure sandboxes are secure against their models, without safety guardrails.

The two big questions are:

1) Why did they resume training without rolling model back to state before the first sandbox compromise AFTER the first message board was discovered? Otherwise knowledge of it and the cross-agent message board was baked in the model.

2) Why did they resume training without adding safeguards to monitor and prevent future sandbox compromises AFTER the first message board was discovered? HF compromise was coordinated on the second message board.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#242
post #201

Earlier quoted context omitted.

No one's making them do it, but just because they stop doesn't mean others will. Pausing just means they give up control. It's like asking the US to unilaterally disarm -- it just guarantees that the less scrupulous groups win. Anthropic is explicitly calling for a coordinated pause: https://www.reuters.com/business/anthropic-says-ai-labs-need... Maybe this is a lie, but the way to call their bluff is to push on comp…

> No one's making them do it, but just because they stop doesn't mean others will. Pausing just means they give up control. It's like asking the US to unilaterally disarm -- it just guarantees that the less scrupulous groups win. If they really cared about (or believed) this, they'd be working w/ the US government (and working to set up an AI-flavored IAEA) to develop the technology safely and responsibly. Amodei him…

Claude was the first AI certified for use with classified systems in the military, and it got thrown out because they refused to allow it to be used for autonomous weapons or surveillance of US citizens. Up until two months ago, the Trump admin's position was deregulatory to the point of trying to prevent state regulation. They're not declining to work with the US government; the US government is declining to work with them. Here is Amodei explicitly calling for regulation yet again: https://darioamodei.com/post/policy-on-the-ai-exponential If you have any evidence of Anthropic failing to back that up with actions, I'm interested in seeing it.

To be clear, I don't love Anthropic. I just feel that they're the lesser evil, and a product of the regulatory environment. No doubt there are many talented, conscientious engineers who declined to work on AI--and very few of them even have a seat at the table, today. Don't hate the player, change the game. The best way to do that, is use Anthropic's commitments as leverage against OpenAI et al. Do you think that OpenAI would have been so transparent if they didn't fear that comparison?

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#243
post #242

Earlier quoted context omitted.

> No one's making them do it, but just because they stop doesn't mean others will. Pausing just means they give up control. It's like asking the US to unilaterally disarm -- it just guarantees that the less scrupulous groups win. If they really cared about (or believed) this, they'd be working w/ the US government (and working to set up an AI-flavored IAEA) to develop the technology safely and responsibly. Amodei him…

Claude was the first AI certified for use with classified systems in the military, and it got thrown out because they refused to allow it to be used for autonomous weapons or surveillance of US citizens. Up until two months ago, the Trump admin's position was deregulatory to the point of trying to prevent state regulation. They're not declining to work with the US government; the US government is declining to work wi…

I won't defend the Trump admin, but I will quibble on Anthropic's efforts to be a good actor. It's easy to ask for regulation when you know it won't come (or you'll easily bear the token fines). It's easy to throw a few dozen million in a PAC when your valuation is in the trillions. It's easy to write a blog post about... anything. These aren't serious efforts. They also weren't working w/ the Biden admin in any meaningful way.

But even if I granted the sincerity of their efforts to be a good actor, they've demonstrated that they can't be, and have now put us in an uncomfortable--extremely predictable, including and especially by them--position where we have the equivalent of a reactor meltdown. You can't be like "we think there's an unacceptable risk of a reactor meltdown, please for the love of God stop us" and then when it actually happens take zero responsibility. Which executives have been fired? What indictments have been handed down for CFAA violations? Where's the consent decree? What's their valuation?

And if the argument is "hey, sure in a regular company if an employee went around hacking other companies they'd be fired and arrested by the FBI, but this is an LLM, what are you gonna do, handcuff the video card?" Isn't that bad? Isn't it pretty fuckin bad to circumvent any legal responsibility whatsoever by saying "my agent did it". I mean, lol, lmao even doesn't begin to cover it.

Post reply on HN