Live data from Hacker News

METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

thezvi.wordpress.com

161–170 of 243 posts

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#161
post #37

Earlier quoted context omitted.

Create a permanent underclass that is unable to access intelligent machines. That’s remarkably dystopian of you.

The alternative is to create a permanent overclass that can hack anyone consequence-free, because they can blame it on AI agents. That also is rather dystopian. Faced with those alternatives, I want neither. Is there a way for us to get neither?

With cybersecurity, it might be "defense dominant" in the sense that we can eventually patch all of our systems to be robust to hacking from even the strongest AI agents. Although it may get worse before it gets better. In a defense dominant world, widespread access to powerful AI could be fine.

However, other areas of risk such as biosecurity may be "offense dominant". For example, we cannot exactly patch the human immune system to defend against artificial viruses the same way that we can patch computer systems.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#162

Earlier quoted context omitted.

I think there's two factors that are worth considering when it comes to this: First, there's an element of timeliness that simply has hard constraints. In order to perform a "proper" analysis of this situation (i.e., little to no dependence on AI tools), you'd have to expect a pretty long wait. I know I'd rather have some sort of "initial report" as quickly as possible than to wait a year or two to get a report about…

They point out the time constraint issue explicitly in the article. But I don't understand how it's been addressed? Like we haven't gotten conclusive data any faster either way, so what's the point? How can it be both so important that we need it so quickly, but at the same time have a tolerance for such plausible deniability? It just doesn't really make sense that both those things are true at the same time.

It would've been great if METR was given more time to conduct their investigation. However, they are an independent organization, and OpenAI only agreed to give them on-premises access for 6 days.

Perhaps if the government decides to sue OpenAI, we could get a more thorough investigation.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#163
post #153

Earlier quoted context omitted.

One could use his flip flop to invalidate his new position. But one of those two positions is true. And if they guy who spent the most energy on position 1 changes his mind, that is worth paying attention to. His second position is likely a more informed, and true, position.

That's totally true, I have no issues with a changing opinion over time. The problem is expressing both opinions with 100% certainty and not updating priors to consider that your new position could also be completely wrong.

Does anyone not admit they could be wrong. All of these people are constantly talking about probability, not saying they are 100% certianty.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#164
post #141

Earlier quoted context omitted.

I mean if doesn't help that Eliezer Yudkowsky spent the first decade of his public life zealously spreading the gospel of the AI singularity as the solution to all mankind's problems. And then later did a 180 degree pivot to AI singularity as the apocalypse with equal zeal and certainty.

One could use his flip flop to invalidate his new position. But one of those two positions is true. And if they guy who spent the most energy on position 1 changes his mind, that is worth paying attention to. His second position is likely a more informed, and true, position.

I don’t see why one of these two positions must be true. And just because he spent a lot of energy trying to convince others of his obsession doesn’t really convince me he has greater insight the future or how the complex consequences unfold.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#165

To clarify, is the TLDR that state of the art models were prompted to cheat / exploit their environment and they did so successfully? Or did OpenAI prompt the models to not cheat and they did anyway? Surprisingly hard to get a clear summary on the basic context of this “incident” separate from marketing lingo and clickbait.

Yeah this feels like crop circles to me. Someone set the context up with an idea in order to catalyze this.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#166
post #97

>1. Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this. I wonder if some of the failures were due to an acquired immunity to "Holy #%^@" moments due to repeated exposure. Like, if you see agents doing surprising things on a regular basis, maybe you…

This is the one.

Every major AI lab is knee deep in weird and mildly demented AIs. They've been dealing with wacky AI shenanigans for so long they've come to expect wacky AI shenanigans. The deviation has been normalized.

It took a high profile "AI oopsie" that went external for OpenAI to lock the fuck in - and take a long look at just how much are their AIs getting up to, and getting away with. I'm still not sure if the lesson would stick.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#167
post #115

I think you have to believe one of two things here. 1. Frontier labs are incapable--either technologically or culturally--of safely developing these powerful systems and should either stop or be forced to stop. At least the FBI should be asking some serious questions (do we really think this is the last time this will happen, at what point are OpenAI complicit, etc) 2. The fuckin thing got out of the cage and all it…

Anthropic has been asking for stronger regulations forever -- and they kept getting criticized for it right here on HN because people assumed it was an attempt at regulatory capture.

Nah the reason is way simpler and more craven: so people like you will post what you just did. They can stop whenever they want; no one's making them do any of this.

There's two possibilities here. One: they know this tech is crazy and they don't care that they can't contain it. Two: they know this tech is mostly bullshit and they don't care they're perpetrating an insane fraud.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#168

Earlier quoted context omitted.

I've read almost nothing about this, but I bet they put an unreliable LLM or 20 on it.

And? What else could they possibly do? Just make the super LLM first, but only ever use it for monitoring lesser LLMs? How will you have monitored the creation of the super LLM?

They could not build the torment nexus.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#169

A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste…

I feel that the main issue with the rationalist crowd is that they live too much in the space of rationality, intelligence and abstractions, but not enough in reality. This leads to an outlook where everything must, almost axiomatically, be intelligible; reality is subordinated to intelligence; and no matter what is real, intelligence can prevail upon it and bend it to its will. Whereas I would argue reality is actua…

To a degree you might say that they live too much in the space of the abstract and not enough in the concrete, or too much in the general and not enough in the specific.

I think that's a fair criticism a lot of the time. Most of the time we have precedent that can guide our actions well. Trying to reason everything out from first principles can be wasteful navel gazing when we've collectively seen the movie a million times.

Focusing on reality and specifics though is what causes people to say there's no global warming because December is cold.

It is sometimes possible to see a trajectory that has never happened before though and that's when you need the autists.

Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

#170

A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste…

The trouble with this is that nobody else was making predictions about AI pre-transformers. Not many are making predictions about AI even now. Forecasting is a preoccupation of the rationalist crowd, and very few people gave much thought to AI before transformers. So the fact that they guessed right about certain things doesn't necessarily mean that the rest of their worldview is sound. A well-informed person who was inclined to making predictions about the future may have drawn similar conclusions without the sci-fi baggage.
Post reply on HN