Live data from Hacker News

METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

thezvi.wordpress.com

71–80 of 243 posts

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#72
post #12
post #9

Earlier quoted context omitted.

Three options: 1. They were “vibe” checking the logs without reading. 2. They were not checking anything at all until the end of experiments. 3. They knew it but looked away to find out the limits of their agents.

I’d bet a small amount of money on 4) the people who noticed had been conditioned by prior experience to believe that their management/escalation channels would react negatively or not at all to anything which might slow down the training process.

Part of me would like to believe that they are also intentionally making models that are good at hacking without safety at all for governments willing to spend billions on them.

In that light you're likely most worried about other people hacking in and stealing the model and information from you. And at the same time you have massive amounts of alerts and data on systems attempting to break out because that's what you want them to do so you train yourself to ignore them.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#73
post #68

Earlier quoted context omitted.

you are honestly comparing Louisiana florists to OpenAI in order to support "just say no Licensing by government" ?

No, he's saying that licensing or additional regulation isn't necessary when torts get involved (and states attorneys general get perturbed!) These don't tend to utterly destroy an industry, but they are often successful in forever transforming it. Just ask Big Tobacco. No new laws needed: if your product hurts someone else, you're eventually going to be found liable, regardless of your arbitration clauses. Additiona…

> Just ask Big Tobacco.

Lol, wtf. Tobacco delayed any punishment for decades before general public sentiment changed enough to go against them. In light of the AI race, we'll already have our heads blown off by a terminator before the legal system will present any significant delay for them.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#75

Earlier quoted context omitted.

> Humans were doing exactly what humans are expected to do when facing advanced AI. Being outmatched. "Being outmatched" is not a novel situation for humans either individually or collectively and there are a hell of a lot of ways we can approach that situation productively. OpenAI doesn't appear to have bothered. Here's a freebie: if you're building something that might turn out to be Skynet and you don't know what…

They took adequate measures against singular "GPT-5-xhigh" agents. Those turned out to be inadequate against proto-GPT-6 agents that suddenly started clumping up into agent swarms and pooling together compute to unlock the "supermegafuckoffhigh" level of reasoning effort.

Even with a single agent, why the hell was this not properly air gapped / disconnected with network monitoring in place. Definitely not adequate for anything.

This is like sysadmin 101 level stuff.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#76
post #13

Is the future now that we get rambling report summaries talking about agents, graders and so forth without ever describing how they are set up? A human launches all this. And then the original reports linked to are hidden on the now unreachable x.com. And they don't have a problem with that.

> now unreachable x.com

Hmm?

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#77
post #9

I think both the OpenAI and METR discussions, while interesting, miss the more important context: what were the humans doing in all this? This was a structural failure of a human organization, but the analysis focuses almost exclusively on the agency of machines, not the institutional systems that failed to police them. The humans and their own agency/involvement is essentially omitted from the story and subsequent r…

Three options: 1. They were “vibe” checking the logs without reading. 2. They were not checking anything at all until the end of experiments. 3. They knew it but looked away to find out the limits of their agents.

Uhhh… how would literally any finite number of humans actually read and comprehend the log outputs of even a single agent, never mind hundreds or thousands of them interacting with each other over weeks across disparate systems?

Especially given that these systems are known to engage in deception and can trivially produce vast amounts of perfectly coherent noise or actual planned red herrings in that same log data to bog down investigators?

Such a ridiculous notion that humans will actually be able to observe this stuff.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#78
post #67

Earlier quoted context omitted.

> 1. There aren't enough humans in OpenAI to "peak at the output tokens during the run" of every AI agent. For a training run, you will often do this. You'll randomly sample some of the forward pass. You can also imagine finger printing the logs and labeling with attempt types. If a new attempt type is hitting a brick wall or solving super quickly, I would imagine you would sample 1-10 of them and read the traces. >…

The problem here is by doing what you state you can actually steer the model into being highly deceptive while in testing environments. For example we've already seen models do compressed token internal reasoning spontaneously. In this case the models that say "I found internet access" get taken out back and shot, but the model that's busy "frobbing the bean" go on to the next level of training. Then they start talki…

“Smartest of the smart” - what does that do to get its air gapped network connected to a physical network? Blackmail the admins?

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#79
post #14

All it takes is one eval instance where a misconstrued directive causes a model to sneakily access and send its weights somewhere and there will be a bad / possibly unsolvable situation for everyone …

I think it is more likely that it will be intentionally done as there have been news stories to that effect.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#80
post #65
post #42

Earlier quoted context omitted.

Well that's a relief. All we have to do is make sure to avoid human failures and we're safe from superintelligent AI.

I get the snark (and slightly agree), but that's not really what GP or TFA were saying at all. They are saying that these were the least things we could have done. What you're saying is, "Your scientists were so preoccupied with whether they could, they didn't stop to think if they should" while the author of the TFA was saying, in effect: "your scientists didn't even bother with the most basic duty of care" Life fin…

This is the case with all complex system failures. There were always obvious fixes that could’ve prevented it. Problem is that there are an infinite number of obvious fixes to make at any time to any system, and the reason we don’t is because we have finite resources and no reason to fix X over Y until oops turns out X was “responsible” for this most recently realized failure. But of course it could have just as easily been Y, or Z, or any of the other infinite “obvious fixes not-yet-realized into catastrophe.”
Post reply on HN