Live data from Hacker News

Anthropic researcher says more than 10% chance AI "could kill all humans"

cbsnews.com

101–110 of 111 posts

Re: Anthropic researcher says more than 10% chance AI "could kill all humans"

#101
post #97

Earlier quoted context omitted.

How on earth do we get a concrete "more than 10%" prediction if the actions of such a system are truly unknowable?

Same way you get "more than 10%" prediction on "I don't know what moves Stockfish will make when I play against it, but I know I will lose". In fact, I will lose in part because I don't know what moves Stockfish will make when I play against it. In my case this is because I am a bad chess player; however it also works for competent chess players: their losses are due to their inability to predict its next move. OK, a…

I am 100% confident that Stockfish will defeat me.

If somebody said "I am 100% confident that AI will destroy humanity" I'd disagree but I'd at least understand how they arrived at that number. But here, why 10%? Why not 50%? Why not 1%?

Re: Anthropic researcher says more than 10% chance AI "could kill all humans"

#102
post #97

Earlier quoted context omitted.

Same way you get "more than 10%" prediction on "I don't know what moves Stockfish will make when I play against it, but I know I will lose". In fact, I will lose in part because I don't know what moves Stockfish will make when I play against it. In my case this is because I am a bad chess player; however it also works for competent chess players: their losses are due to their inability to predict its next move. OK, a…

I am 100% confident that Stockfish will defeat me. If somebody said "I am 100% confident that AI will destroy humanity" I'd disagree but I'd at least understand how they arrived at that number. But here, why 10%? Why not 50%? Why not 1%?

Because there are much more unknown parameters in the outcome of AI for humanity than in your match against Stockfish.

For example, the timeline upon which AIs get effectively smarter than us is uncertain. Let's say you believe the probability this occurs before we solve the alignment problem is 80%, it doesn't seem too far fetched to think that in this case there is at least a 12.5% chance that AIs coordinate against us in a catastrophic way. Combining these probabilities you get a 10% chance of a catastrophic outcome for humanity.

Note that the numbers are not to be taken at face value, I just wanted to give an example of thought process which could give such a figure.

Re: Anthropic researcher says more than 10% chance AI "could kill all humans"

#103
post #97

Earlier quoted context omitted.

Same way you get "more than 10%" prediction on "I don't know what moves Stockfish will make when I play against it, but I know I will lose". In fact, I will lose in part because I don't know what moves Stockfish will make when I play against it. In my case this is because I am a bad chess player; however it also works for competent chess players: their losses are due to their inability to predict its next move. OK, a…

I am 100% confident that Stockfish will defeat me. If somebody said "I am 100% confident that AI will destroy humanity" I'd disagree but I'd at least understand how they arrived at that number. But here, why 10%? Why not 50%? Why not 1%?

If we were actively trying to make this "win" in the Stockfish sense, it would likely be 99%.

We are trying to make a system that doesn't want to "win" in the sense, but wants to "win" by being helpful, harmless, an honest (or some variation of that).

What odds do you put on us making the "helpful, harmless, an honest" part, bug-free? Or rather, that the bugs will be sufficiently minor as to not kill everyone, given that that we're clearly in the world where people not only use it beyond its competence, but also attempt to maliciously subvert all those efforts to make it "harmless" while keeping the "helpful and honest" parts so they can use it to be dangerous.

Anyone who successfully subverts a "helpful, harmless, an honest" training system then goes and does whatever they wanted with this system; right now when they do so, which is near constantly, it happens with a system of limited competence, so they get it to scam or to hack etc.

The reason I would also pick 10% is that I think the constant abuse and misuse (the latter including simply using a system beyond its competence without malice) means we get an escalating series of disasters, which at some point kill enough people that everyone agrees this is madness and stops.

10% is the chance we blow right through all the warning shots and a sufficiently competent AI is either abused or misused (again, misuse can be without malice), resulting in it having a goal (/prompt) that is effectively to win the Stockfish sense.

Re: Anthropic researcher says more than 10% chance AI "could kill all humans"

#104

Earlier quoted context omitted.

In order to understand the rational argument, one needs to follow closely the latest developments of misaligned AI (I think only few are doing so). The most important readings IMO are the METR analysis of the HuggingFace incident and the AISI report of the Github incident. The basic argument is extremely simple: - AIs can, depending on context, pursue a task with complete disregard for humans/values - In the future,…

The basic argument is extremely simple: - goats can, depending on context, pursue a task with complete disregard for humans/values - In the future, goats will have enormously more means and smarts - A goat could then assess that humans are an impediment to its tasks, escape containment and proceed. You really need to raise goats, you'll be surprised. ------ As far as I can tell, AIs are like smart farm animals. I use…

You clearly haven't read anything about AI accidents and late developments, besides headlines.

Re: Anthropic researcher says more than 10% chance AI "could kill all humans"

#106
post #94

Earlier quoted context omitted.

The probability calculator doesn't really help, since the core question for me is how one arrives at its "Probability that misalignment leads to an unrecoverable global catastrophe.".

Sure, sure. There's many others like this to help you combine whatever you do feel you can put a number to. For me, that particular question is "probably 0, but with 100% variance". This is because I think most of the things AI can do harm with are small enough to force us to take the risk seriously, and only a few are big enough to get us all before we take the risk seriously.

This is a succinct explanation for what I have been thinking as well, thank you for putting it into words.

Re: Anthropic researcher says more than 10% chance AI "could kill all humans"

#107
post #73

> 10% chance AI "could kill all humans" Why not 10% chance that it will create enormous prosperity for all ? This is why the average person is increasing pissed at AI in general. That it gets associated with negativity.

I'm old enough to remember when people dismissed all the doom coming from these companies as "marketing". (A thing many of them have been entirely consistent about since GPT-2, or indeed earlier given the founding documents). I know a few people around these circles; People like this are quite sincere about the risk, and that they think poorly of their bosses and how risk is being handled.

But this is marketing. This is just pre-IPO talk.

Anthropic’s goal is not anything to do with AI, it is purely to generate more profit than any other company.

Securing/lobbying to regulate their competitors through doom talk, or promising that they will bankrupt every other company because AI will “do it all, better than everyone else” is purely for investment purpose.

As long as their incentive is monetary, it’s not unreasonable to believe this is just marketing.

Re: Anthropic researcher says more than 10% chance AI "could kill all humans"

#108

It's all marketing. Don't fall for it, it's the age old strategy "our product is extremely dangerous, so fear us". Like the tobacco companies saying all the time "our cigarettes are so dangerous, they cause cancer". Or like Purdue Pharma saying "do not use fentanyl, it's so dangerous, it kills thousands of people every year". They try to generate fear in their products, to shock investors into buying their stock.

What even is the marketing strategy though? Generate a buzz headline so people engage and speak about the company? Isn't commenting falling for it? Posts here with no activity can't climb gravity without engagement.

“We have built this thing that is so powerful that it will wipe out all jobs ever created. Therefore we are the only thing worth investing in.”

Re: Anthropic researcher says more than 10% chance AI "could kill all humans"

#109
post #70
post #13

I’ve yet to see a rational argument for how we go from super intelligent LLMs to human extinction or extermination. I understand that some smart people are worried about it. I just haven’t come across a believable or understandable argument.

It's not too hard to imagine potential scenarios, some example have been given in previous responses. But there is another kind of argument to be made: if you play chess against a player that is far smarter than you (chess wise), you know you are going to lose, even if you don't know how. So the mere existence of a smarter species than us is a threat in itself.

[deleted]

Re: Anthropic researcher says more than 10% chance AI "could kill all humans"

#110

Earlier quoted context omitted.

climate change was never going to kill all humans.

There's some scenarios where it could. One plausible one off the top of my head is: climate change makes the flow of the Indus (which has its headwaters in the glaciers of the Himalayas) less reliable. Two nuclear powers (India and Pakistan) are highly reliant on the Indus for agriculture. Things escalate out of control, and when nukes start flying it trips the terrifyingly ramshackle Dead Hand system that Russia's g…

yeah i mean i'm gonna chalk that one up to the nukes.
Post reply on HN