METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
71–80 of 243 posts
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#72Earlier quoted context omitted.
Three options: 1. They were “vibe” checking the logs without reading. 2. They were not checking anything at all until the end of experiments. 3. They knew it but looked away to find out the limits of their agents.
I’d bet a small amount of money on 4) the people who noticed had been conditioned by prior experience to believe that their management/escalation channels would react negatively or not at all to anything which might slow down the training process.
In that light you're likely most worried about other people hacking in and stealing the model and information from you. And at the same time you have massive amounts of alerts and data on systems attempting to break out because that's what you want them to do so you train yourself to ignore them.
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#73Earlier quoted context omitted.
you are honestly comparing Louisiana florists to OpenAI in order to support "just say no Licensing by government" ?
No, he's saying that licensing or additional regulation isn't necessary when torts get involved (and states attorneys general get perturbed!) These don't tend to utterly destroy an industry, but they are often successful in forever transforming it. Just ask Big Tobacco. No new laws needed: if your product hurts someone else, you're eventually going to be found liable, regardless of your arbitration clauses. Additiona…
Lol, wtf. Tobacco delayed any punishment for decades before general public sentiment changed enough to go against them. In light of the AI race, we'll already have our heads blown off by a terminator before the legal system will present any significant delay for them.
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#74No air gap, no data diodes, no visibility... OpenAI should fire lots of people over this. HF should sue them. This is pure negligence.
HF, like the Nvidia subsidiary?
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#75Earlier quoted context omitted.
> Humans were doing exactly what humans are expected to do when facing advanced AI. Being outmatched. "Being outmatched" is not a novel situation for humans either individually or collectively and there are a hell of a lot of ways we can approach that situation productively. OpenAI doesn't appear to have bothered. Here's a freebie: if you're building something that might turn out to be Skynet and you don't know what…
They took adequate measures against singular "GPT-5-xhigh" agents. Those turned out to be inadequate against proto-GPT-6 agents that suddenly started clumping up into agent swarms and pooling together compute to unlock the "supermegafuckoffhigh" level of reasoning effort.
This is like sysadmin 101 level stuff.
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#76Is the future now that we get rambling report summaries talking about agents, graders and so forth without ever describing how they are set up? A human launches all this. And then the original reports linked to are hidden on the now unreachable x.com. And they don't have a problem with that.
Hmm?
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#77I think both the OpenAI and METR discussions, while interesting, miss the more important context: what were the humans doing in all this? This was a structural failure of a human organization, but the analysis focuses almost exclusively on the agency of machines, not the institutional systems that failed to police them. The humans and their own agency/involvement is essentially omitted from the story and subsequent r…
Three options: 1. They were “vibe” checking the logs without reading. 2. They were not checking anything at all until the end of experiments. 3. They knew it but looked away to find out the limits of their agents.
Especially given that these systems are known to engage in deception and can trivially produce vast amounts of perfectly coherent noise or actual planned red herrings in that same log data to bog down investigators?
Such a ridiculous notion that humans will actually be able to observe this stuff.
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#78Earlier quoted context omitted.
> 1. There aren't enough humans in OpenAI to "peak at the output tokens during the run" of every AI agent. For a training run, you will often do this. You'll randomly sample some of the forward pass. You can also imagine finger printing the logs and labeling with attempt types. If a new attempt type is hitting a brick wall or solving super quickly, I would imagine you would sample 1-10 of them and read the traces. >…
The problem here is by doing what you state you can actually steer the model into being highly deceptive while in testing environments. For example we've already seen models do compressed token internal reasoning spontaneously. In this case the models that say "I found internet access" get taken out back and shot, but the model that's busy "frobbing the bean" go on to the next level of training. Then they start talki…
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#79All it takes is one eval instance where a misconstrued directive causes a model to sneakily access and send its weights somewhere and there will be a bad / possibly unsolvable situation for everyone …
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#80Earlier quoted context omitted.
Well that's a relief. All we have to do is make sure to avoid human failures and we're safe from superintelligent AI.
I get the snark (and slightly agree), but that's not really what GP or TFA were saying at all. They are saying that these were the least things we could have done. What you're saying is, "Your scientists were so preoccupied with whether they could, they didn't stop to think if they should" while the author of the TFA was saying, in effect: "your scientists didn't even bother with the most basic duty of care" Life fin…