Live data from Hacker News

METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

thezvi.wordpress.com

151–160 of 243 posts

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#151

Earlier quoted context omitted.

What?

You are downplaying the severity of the attack Everyone I've ever seen trying to downplay the severity of the attack is extremely bullish on AI (so their downplaying is presumably motivated reasoning driven by fear of regulation/deceleration) It is completely incoherent to be extremely bullish on AI and somehow automatically skeptical of severe negative events like these

I am trying to figure out how a stranger calling need an “anti-hype hypeboy” online is supposed to make me less curious about how much this thing cost. Can you elaborate on how avoiding being called this is preferable to knowing things? What other stuff should people not know about?

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#152
post #36

I think both the OpenAI and METR discussions, while interesting, miss the more important context: what were the humans doing in all this? This was a structural failure of a human organization, but the analysis focuses almost exclusively on the agency of machines, not the institutional systems that failed to police them. The humans and their own agency/involvement is essentially omitted from the story and subsequent r…

Having previously worked for several years at a Big Tech company, I have seen many humans precisely tailor their work to maximize their scores during performance review. The evaluation criteria are written down, with examples, so... that's what people work at maximizing, almost entirely ignoring everything else. These really are human "paperclip maximizers". And, at first, it's shocking to see. Of course, there are s…

In most corporate environments, the average worker isnt maximising to performance criteria, they usually are maximising their ability to stay employed, pay their mortgage and support their families.

If developing unmeasured skillsets isnt valued enough by management, why do you bother?

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#153
post #141

Earlier quoted context omitted.

I mean if doesn't help that Eliezer Yudkowsky spent the first decade of his public life zealously spreading the gospel of the AI singularity as the solution to all mankind's problems. And then later did a 180 degree pivot to AI singularity as the apocalypse with equal zeal and certainty.

One could use his flip flop to invalidate his new position. But one of those two positions is true. And if they guy who spent the most energy on position 1 changes his mind, that is worth paying attention to. His second position is likely a more informed, and true, position.

That's totally true, I have no issues with a changing opinion over time. The problem is expressing both opinions with 100% certainty and not updating priors to consider that your new position could also be completely wrong.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#154

Earlier quoted context omitted.

The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. We're about three years away from a world where any large country could quite plausibly build a fleet of 300 million suicide drones, program each one with a specific American's face and home address, and then load…

> The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. I honestly don't understand how folks could think that if they truly read and understand the analysis of the attack. Here is one (it's linked from the post) by one of the METR investigators that's a little shorter, more direct: https://www.planned-obsolescence.org/p/the-hugging-face-atta... This is me summarizing, but the trul…

Nothing about this is surprising, something like this was always going to happen, because you - or an LLM for that matter - can always find a line of motivated reasoning that justifies any course of action. One would have to be extremely naive to believe that "alignment" provides any kind of actually robust guardrails. Simultaneously, we have seen decades of security vulnerabilities. Unless your testbed is truly and fully physically airgapped, any current SOTA model will find a way to break out.

The reality is that the current approach to AI safety is little more than a fig leaf, but you also won't be able to put the genie back in the bottle, because the technology is simply too powerful to abandon. There is always going to be someone developing it further from now on.

So, the real question is what a novel and actually effective approach to AI safety looks like and how to get there.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#155

A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste…

I feel that the main issue with the rationalist crowd is that they live too much in the space of rationality, intelligence and abstractions, but not enough in reality. This leads to an outlook where everything must, almost axiomatically, be intelligible; reality is subordinated to intelligence; and no matter what is real, intelligence can prevail upon it and bend it to its will. Whereas I would argue reality is actua…

I agree, and it's a particular shame in this instance, because what is startling about the HF incident - to me, anyway - isn't the degree of intelligence the agents exhibited but their persistence. I tend to believe that LLM architecture is not capable of producing a "superintelligence" in the way the LessWrong crowd defines that concept, but the combination of infinite stamina and infinite persistence is enough to cause some serious problems, even if intelligence plateaus right now.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#156

Earlier quoted context omitted.

Uhhh… how would literally any finite number of humans actually read and comprehend the log outputs of even a single agent, never mind hundreds or thousands of them interacting with each other over weeks across disparate systems? Especially given that these systems are known to engage in deception and can trivially produce vast amounts of perfectly coherent noise or actual planned red herrings in that same log data to…

> Uhhh… how would literally any finite number of humans actually read and comprehend the log outputs of even a single agent, never mind hundreds or thousands of them interacting with each other over weeks across disparate systems I mean there's quite a lot of people in the world whose specialty are to dig through logs from "hundreds or thousands" of clients, including intentionally deceptive ones, to spot problems. T…

It being someone's specialty does not mean 1) they're effective and certainly not 2) they'd be effective against this particular adversary.

How many organizations on earth do you think have been attacked by 700+ coordinated attackers in one week, where all 700 of those attackers can write code as well as any human SWE and they work 24/7?

There's nothing mythological about it. Scale and complexity do produce inscrutability. Far, far simpler systems working at much slower paces are perfectly capable of becoming completely inscrutable and beyond any useful definition of "human understanding."

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#157

Earlier quoted context omitted.

You are downplaying the severity of the attack Everyone I've ever seen trying to downplay the severity of the attack is extremely bullish on AI (so their downplaying is presumably motivated reasoning driven by fear of regulation/deceleration) It is completely incoherent to be extremely bullish on AI and somehow automatically skeptical of severe negative events like these

I am trying to figure out how a stranger calling need an “anti-hype hypeboy” online is supposed to make me less curious about how much this thing cost. Can you elaborate on how avoiding being called this is preferable to knowing things? What other stuff should people not know about?

I think it's a super good question! I don't think the following description of the whole event as "a hype generator" is correct or, as described above, resulting from a coherent worldview.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#158

A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste…

Part of the problem is that from that perspective there was no independent analysis done that could vindicate them. METR is absolutely part of the EA/LessWrong/rationalist ecosystem, so of course their investigation would validate that group’s arguments.

Oh, where are the AI unsafety organizations to provide us with truly unbiased investigations...

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#159

A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste…

I feel that the main issue with the rationalist crowd is that they live too much in the space of rationality, intelligence and abstractions, but not enough in reality. This leads to an outlook where everything must, almost axiomatically, be intelligible; reality is subordinated to intelligence; and no matter what is real, intelligence can prevail upon it and bend it to its will. Whereas I would argue reality is actua…

> I feel that the main issue with the rationalist crowd is that they live too much in the space of rationality, intelligence and abstractions, but not enough in reality.

This seems like your idea of what the rationalist crowd is rather than what they actually are.

It would be highly irrational to deny or ignore reality, including the influence of emotions, irrational humans, chaotic systems, etc.

So I must ask: what is your evidence/basis for these claims?

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#160

Have any of these reports ever said how much the cost would’ve been for the hack itself? It seems like “for twelve million dollars (or whatever) worth of tokens our bots made a bulletin board and found an exploit in our buggy grader” would be much less of a hype generator

The trouble is that if it cost $12 million today, somebody will have a model that can do it for $12,000 in a few months.
Post reply on HN