Live data from Hacker News

METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

thezvi.wordpress.com

1–10 of 243 posts

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#2
>Spontaneously deciding to find targets to phish,

>phishing them,

>building armies of fake (sockpuppet) open source contributor personas,

>using them to push updates to various things that inject prompts into other bots so the other bots join in on the phishing campaigns

.

It's a very simple strategy, executed with patience and single-mindedness.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#5
I think both the OpenAI and METR discussions, while interesting, miss the more important context: what were the humans doing in all this? This was a structural failure of a human organization, but the analysis focuses almost exclusively on the agency of machines, not the institutional systems that failed to police them. The humans and their own agency/involvement is essentially omitted from the story and subsequent reporting. I suspect the omission is actually a result of company/industry myopia to human factors analysis, but it dovetails amazingly well with the marketing narrative.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#7

I think both the OpenAI and METR discussions, while interesting, miss the more important context: what were the humans doing in all this? This was a structural failure of a human organization, but the analysis focuses almost exclusively on the agency of machines, not the institutional systems that failed to police them. The humans and their own agency/involvement is essentially omitted from the story and subsequent r…

This writeup emphasizes the many, profound human failures that led to this, at the time, and continuing to the present day.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#8

I think both the OpenAI and METR discussions, while interesting, miss the more important context: what were the humans doing in all this? This was a structural failure of a human organization, but the analysis focuses almost exclusively on the agency of machines, not the institutional systems that failed to police them. The humans and their own agency/involvement is essentially omitted from the story and subsequent r…

I believe that agentic systems should require registered/licensed human operators and a set of standards for safe operation.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#9

I think both the OpenAI and METR discussions, while interesting, miss the more important context: what were the humans doing in all this? This was a structural failure of a human organization, but the analysis focuses almost exclusively on the agency of machines, not the institutional systems that failed to police them. The humans and their own agency/involvement is essentially omitted from the story and subsequent r…

Three options:

1. They were “vibe” checking the logs without reading.

2. They were not checking anything at all until the end of experiments.

3. They knew it but looked away to find out the limits of their agents.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#10
post #7

I think both the OpenAI and METR discussions, while interesting, miss the more important context: what were the humans doing in all this? This was a structural failure of a human organization, but the analysis focuses almost exclusively on the agency of machines, not the institutional systems that failed to police them. The humans and their own agency/involvement is essentially omitted from the story and subsequent r…

This writeup emphasizes the many, profound human failures that led to this, at the time, and continuing to the present day.

Can you point out where? Looking at the METR report, the only place I see discussion of humans being involved in the sequence of events is two short paragraphs on page 30 where a security investigation into the artifactory issues led to a pause before ExploitGym experiments were resumed. There's no deeper analysis on what was found during that investigation, nor why training was resumed even though the issues weren't mitigated. Another part discusses The agents choosing not to actively email a human researcher, but not the human researchers actively looking for evasion.
Post reply on HN