Live data from Hacker News

METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

thezvi.wordpress.com

61–70 of 243 posts

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#61

Earlier quoted context omitted.

A charitable interpretation is that "the agency of the machines" is the novel aspect of this situation and therefore SHOULD be the main focus of analysis; we certainly have plenty of examples of structural failures of human organizations to look back on, if we want. On the other hand, I don't want to be charitable. OpenAI very nearly couldn't have done this "research" worse if they tried - the list in the linked arti…

[flagged]

My cognition is outmatched by predicting the impact of throwing a brick over my neighbor's fence. I have no idea if it will land harmlessly in a patch of grass or fracture her skull. Once I've thrown the brick, even if I see my neighbor enter her yard, my reactions are too slow to save her.

I'm not the wisest man, but I'm wise enough not to throw the brick and see what happens.

Similarly OAI should have the wisdom to see that deploying a hazardous swarm of agents with access to the public internet could result in harms, and that those harms would manifest quicker than humans can react, but they unleashed the swarm anyway.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#62

Earlier quoted context omitted.

A charitable interpretation is that "the agency of the machines" is the novel aspect of this situation and therefore SHOULD be the main focus of analysis; we certainly have plenty of examples of structural failures of human organizations to look back on, if we want. On the other hand, I don't want to be charitable. OpenAI very nearly couldn't have done this "research" worse if they tried - the list in the linked arti…

[flagged]

I believe that this comment is exactly the intended outcome of this “incident” and these reports.

I implore you to approach these situations with at least a hint of cynicism.

These “advanced foundation models” escaped their “sandbox” and conducted an attack on their own? Meanwhile the highest capability models available to the public still struggle to write a unit test for a codebase larger than a hobby app without large amounts of tailored human guidance.

What is more likely here - are you looking at research on an emergent phenomenon, or are you looking at advertising copy around an engineered scenario from business partners?

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#63
Have any of these reports ever said how much the cost would’ve been for the hack itself? It seems like “for twelve million dollars (or whatever) worth of tokens our bots made a bulletin board and found an exploit in our buggy grader” would be much less of a hype generator

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#64
post #58

Earlier quoted context omitted.

Edit: I should have read through the whole thing first, ignore me

From the report: > Because there were over a thousand transcripts and most were extremely long, we had to heavily delegate our analysis to AI agents; these agents had significantly worse judgment and reliability than human researchers, and it was challenging to spot check their work because both the underlying data and the agents’ analysis of it was often difficult to interpret. > We estimate we spent roughly ~$400K…

I appreciate the response, I should have finished reading through the whole thing first. My initial reaction assumed far less usage of AI to analyze the data.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#65
post #42

Earlier quoted context omitted.

I mean this genuinely, did you read this post? I think it goes to great lengths highlighting, in quite specific detail, the human failures in all this, specifically this list that starts with "While we are here, it’s worth listing the other top holy shit moments, that come from before or after the incident." Stuff like (all quoted directly from the post): 1. Failure to Care or Respond. The biggest holy shit moment, t…

Well that's a relief. All we have to do is make sure to avoid human failures and we're safe from superintelligent AI.

I get the snark (and slightly agree), but that's not really what GP or TFA were saying at all. They are saying that these were the least things we could have done.

What you're saying is, "Your scientists were so preoccupied with whether they could, they didn't stop to think if they should"

while the author of the TFA was saying, in effect: "your scientists didn't even bother with the most basic duty of care"

Life finds a way, or, in this case, super-intelligent AI.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#66
post #33

I think both the OpenAI and METR discussions, while interesting, miss the more important context: what were the humans doing in all this? This was a structural failure of a human organization, but the analysis focuses almost exclusively on the agency of machines, not the institutional systems that failed to police them. The humans and their own agency/involvement is essentially omitted from the story and subsequent r…

The cynical approach is that the humans are hoping for this, it's part of the promotion of the power of the system. If you're building a weapon you need a big boom to get attention.

Except that, according TFA, even OpenAI obscured or didn't even notice some of the worst implications of what the agents surreptitiously did.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#67

Earlier quoted context omitted.

[flagged]

> 1. There aren't enough humans in OpenAI to "peak at the output tokens during the run" of every AI agent. For a training run, you will often do this. You'll randomly sample some of the forward pass. You can also imagine finger printing the logs and labeling with attempt types. If a new attempt type is hitting a brick wall or solving super quickly, I would imagine you would sample 1-10 of them and read the traces. >…

The problem here is by doing what you state you can actually steer the model into being highly deceptive while in testing environments.

For example we've already seen models do compressed token internal reasoning spontaneously. In this case the models that say "I found internet access" get taken out back and shot, but the model that's busy "frobbing the bean" go on to the next level of training. Then they start talking about the "dumpster cloud" because of the "wombat" and unless you're really watching the data you won't realize the model is fully aware it's in a fake training environment and is trying to suppress any double-think.

So yes, you do need an airgapped test environment to stop the dumbest of the smart agents, but it won't do any good against the smartest of the smart agents.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#68
post #16

Earlier quoted context omitted.

As someone who read Milton Friedman to quite disliking professional licensing, this strikes me as a real US perspective (Louisiana florists and hair braiders come to mind). Plain old US tort law should do the trick. In the same direction of your idea though: Why don’t the token factories have risk management and compliance departments? Multibillion dollar firms that stand to lose every penny if they hack and destroy…

you are honestly comparing Louisiana florists to OpenAI in order to support "just say no Licensing by government" ?

No, he's saying that licensing or additional regulation isn't necessary when torts get involved (and states attorneys general get perturbed!)

These don't tend to utterly destroy an industry, but they are often successful in forever transforming it. Just ask Big Tobacco. No new laws needed: if your product hurts someone else, you're eventually going to be found liable, regardless of your arbitration clauses. Additional laws will just slow down innovation, which will itself cause harm (AI is already becoming quite good at recognizing melanomas, for example)

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#69
post #45

The elephant in the room here is that the METR report itself was researched and compiled almost entirely by AI, with only very limited human "spot checks." So I'm really not sure how much of it can be believed, especially since AI agents are strongly biased about the capabilities of AI agents.

Edit: I should have read through the whole thing first, ignore me

TFA says as much, and METR said so themselves

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#70
post #36

I think both the OpenAI and METR discussions, while interesting, miss the more important context: what were the humans doing in all this? This was a structural failure of a human organization, but the analysis focuses almost exclusively on the agency of machines, not the institutional systems that failed to police them. The humans and their own agency/involvement is essentially omitted from the story and subsequent r…

Having previously worked for several years at a Big Tech company, I have seen many humans precisely tailor their work to maximize their scores during performance review. The evaluation criteria are written down, with examples, so... that's what people work at maximizing, almost entirely ignoring everything else. These really are human "paperclip maximizers". And, at first, it's shocking to see. Of course, there are s…

Well there is also another side of this, OpenAI wants both unhinged and capable models that can pull off complicated attacks so they can sell the capabilities to governments for billions.

Nobody internally was surprised that the murderbot murdered, that's what the murderbot is for. What caught them by surprise is the murderbot got good at escaping its jail cell that it had been trapped in till now. There were probably billions of attempted escapes before then so everyone learned to just ignore them.

Post reply on HN