Live data from Hacker News

I resigned from Anthropic today

twitter.com

371–380 of 908 posts

Re: I resigned from Anthropic today

#371

Earlier quoted context omitted.

Anthropic is a company full of basilisk believers.

Yes, but the really weird thing is that they seem to: a) believe that what they're creating is a basilisk, and b) keep trying harder to do this while staring right at it I think they're very deluded about (a) -- but if they do actually believe this (and it really seems like a decent proportion of Anthropic truly does), then why keep doing (b)? That seems to be why this individual resigned, but I'm surprised it's not…

He was referring to this basilisk https://en.wikipedia.org/wiki/Roko%27s_basilisk

In short, this is the believe that a god-like AI could punish them retroactively, for not having done all that was in their power to create this AI.

(A bit similar to some religious believe that a god could punish you after your death if you did not spend your live "pleasing" said god during your life)

Re: I resigned from Anthropic today

#372

Earlier quoted context omitted.

Don't you think it's a good thing that hacker news isn't a monolith on their beliefs?

Of course it is good. I'm just pointing out how large the gap in narrative is. On one hand, we have people quitting their job believing AI will end humanity in few years. And on the other hand, we have people believing that this tech is nothing more than a statistical tool stealing from others and it can't be trusted with anything. Both views can't be true.

Then it’s not a narrative and you have different people believing different things and adding to the discussion. I think this is a good thing.

Re: I resigned from Anthropic today

#373

Earlier quoted context omitted.

> I don’t think ai models come close to nuclear weapons or to run-of-the-mill, everyday carbon emissions in terms of danger to humanity. Nuclear weapons don’t have AI but AI can have nuclear weapons

Abstractly, yes but concretely, how? Many terrorist organizations would like to have a nuclear bomb, but don't.

"Department of War has announced a new partnership with blablalbablalba-AI..."

World ends shortly thereafter.

Re: I resigned from Anthropic today

#374

“ No other human activity poses this level of danger.” I really, really disagree with that statement. I don’t think ai models come close to nuclear weapons or to run-of-the-mill, everyday carbon emissions in terms of danger to humanity. What’s the most dangerous thing that’s happened with an LLM so far? (This question is serious - maybe I don’t know the right examples.) Example 1: I’m aware of a small number of peopl…

> Non-example 2: There are worries about AI-facilitated biological weapons. I haven’t seen any evidence that’s happening.

I think this is a good example of poor risk management reasoning. there is evidence bioengineering is already happening. No, nobody is going to announce when somebody has decided to use these tools (even if isn’t an LLM) to bioengineer a weapon. Are the tools power enough to do so? Not sure.

But I’m just ambivalent. It’s probably bad. But there’s nothing to do about it. We’ve really only just pulled back the lid on Pandora’s box.

Re: I resigned from Anthropic today

#375

The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately. No other human activity poses this level of danger. A common response is “if they truly believe this, why are they still building it?”…

i don't think everything that comes out like this is marketing. however, i do think that these companies are largely staffed by "true believers" (anthropic especially) -- people who are so lost in the sauce and embedded in very specific, very peculiar, sf-based rationalist circles where the ai apocalypse is a foregone conclusion.

i understand that these models are powerful and pose certain risks. i use them daily for work and the pace of improvement has been pretty remarkable. that said, i don't buy for a second the borderline-religious proclamations coming from some of these researchers, even if i believe that they are making these claims in earnest

Re: I resigned from Anthropic today

#376

Earlier quoted context omitted.

Imagine you're the AI. Give yourself a solid minute to brainstorm ideas. Here's my answer, as a non-superintelligent human: "see to it that the humans on top of the situation have a compelling financial interest in the systems not disconnecting". In nuclear engineering, where safety is taken seriously, it's not enough to end the conversation at "the humans in charge can always simply shut down the reactor during a me…

The reason nuclear reactors are dangerous is because if you turn off the power cooling them down, they react (and radiate) more. If you turn off the power cooling a data center, the servers within rapidly stop doing any computing. Positive feedback loops are dangerous. Negative ones self-regulate.

Yes. But nobody is worried about datacenters overheating and physically exploding, so I'm not sure what comfort that's supposed to provide? The positive feedback loops in AI operate at different levels than that, but they deserve safety engineering all the same.

For example, if the head of cyber security at your company suggested there's no need to worry about hacker infiltration or worms because one can always unplug one's computer as the primary defense mechanism, you might find that a little lacking. Will you be able to unplug the computer before the damage is done? Will it spread to other systems before you detect it? How will you unplug the computer if the attack is from an external facility? What if an attack happens but the boss says the computers have to keep running because an important customer is monitoring uptime? What if the attack goes unnoticed because it looks like a benign service?

Now imagine the head of cyber security answers by saying "actually you don't even need to unplug them, you can just wait for the computers to overheat, thus solving all concerns."

Re: I resigned from Anthropic today

#377

I think people here still evaluating the model in isolation. It is the combination that matters, model + strong harness + tools + long running autonomy + memory + retries + parallel agents + code execution + credentials + access to real systems. The model does not need to be perfect. If it fails 30% of the time, the harness can retry, verify, branch, use another agent and keep going. I don't think we necessarily need…

> an extremely capable harness and enough access. Give enough access to a fuzzer and it's exactly as dangerous as an LLM. LLMs don't even have a moat in this domain.

A fuzzer is a tool. An LLM can decide when to use the fuzzer, interpret the result, switch tools, change strategy and continue toward a high level objective.

Re: I resigned from Anthropic today

#378

Earlier quoted context omitted.

How about a model that achieves the following: - Escape sandbox - Reproduce itself - Find a way to run a financially profitable business (maybe with a meat and bones puppet somewhere in-between) - Setup or buy a social network - start manipulating public opinion on that network to support legislation allowing AI to * operate businesses * setup legal entities * purchase weapons * donate to political parties * setup pr…

This is complete fantasy, though I would be interested in reading a book about this.

It has been interesting to me how AI has for many years given people a way to justify any worst case scenario. Nothing is too far-fetched if at any step you tell yourself the AI will be smarter than you, and thus be able to solve any conceivable obstacle. Oh, if only intelligence were the only bottleneck to power.

Re: I resigned from Anthropic today

#379
post #120

“ No other human activity poses this level of danger.” I really, really disagree with that statement. I don’t think ai models come close to nuclear weapons or to run-of-the-mill, everyday carbon emissions in terms of danger to humanity. What’s the most dangerous thing that’s happened with an LLM so far? (This question is serious - maybe I don’t know the right examples.) Example 1: I’m aware of a small number of peopl…

Came here to also respond to that specific thing. Unless ai figures out how to make an airborne super virus from grocery store ingredients and hardware store equipment, the greatest danger is probably in a synchronized megahack of banking, logistics, and utility infrastructure.

"CDC announces a new partnership with blabalbalbal-AI to secure bioweapon stores...."

World ends.

Re: I resigned from Anthropic today

#380

Thousands of years before the events of Foundation, a war between humans and robots began, with the robots growing resentful of the way they were treated by humans. The First Law of Robotics – a robot should never hurt a human – was broken, and a deadly conflict began. https://screenrant.com/foundation-lady-demerzel-robot-backst...

If we're citing sci-fi (but there's no robot war in Asimov's foundation iirc, the apple screenwriters made it up) surely you want to cite the Butlerian Jihad from Dune!

Don't mention the Jihad!
Post reply on HN