Live data from Hacker News

We could stumble into AI catastrophe

cold-takes.com

121–122 of 122 posts

Re: We could stumble into AI catastrophe

#121

Earlier quoted context omitted.

Sure, that is a risk. I'm just saying that I'd rather wait until that point where we actually see AIs making poor attempts to estimate their operators' knowledge and deceive them; then, I'd have no problem worrying about it. But I suspect that point may not come for a long while, so until then it remains speculation.

You may enjoy this after-action report of a person being attacked by a hostile AI, played by ChatGPT: https://www.lesswrong.com/posts/9kQFure4hdDmRBNdH/how-it-fee... Now to be clear, this is possibly the easiest mark conceivable, and the poster played into his own demise at any available opportunity. But we should expect the first marks to be easy marks. People looking at videos of Hitler today don't understand why a…

Good point. I've thought for a while now that one of the big risks of present AI chatbots is that they can endlessly validate any arbitrary thoughts of an emotionally unstable human operator, creating a feedback loop that results in a strong irrational conviction, that might even be harmful to the operator or to others. (See: Lemoine and LaMBDA, or the person you linked to.)

So I guess the risk there would be, the AI is accessible by the public, some fraction of the public "radicalizes" itself in a similar direction from talking with the AI too much, these people form a coherent movement, and that movement becomes powerful enough to take over the world. So then the question becomes, just how plausible is it for grassroots rebellions so formed to succeed against the authorities in real life, as opposed to fiction?

The fundamental issue here is that the attack can occur through untargeted manipulation of vulnerable people, as opposed to the targeted manipulation of specific people in power (which I suspect near-future AI models will still be incapable of). The obvious defense would be to only allow AIs to be prompted by a sufficiently-large committee, alongside some social machinery in place to make sure the committee isn't all colluding as part of a cult. But I doubt many of the AI-risk people would ever accept that, operating under the shockingly common "powerful aligned AI or bust" model.

Re: We could stumble into AI catastrophe

#122

Earlier quoted context omitted.

You may enjoy this after-action report of a person being attacked by a hostile AI, played by ChatGPT: https://www.lesswrong.com/posts/9kQFure4hdDmRBNdH/how-it-fee... Now to be clear, this is possibly the easiest mark conceivable, and the poster played into his own demise at any available opportunity. But we should expect the first marks to be easy marks. People looking at videos of Hitler today don't understand why a…

Good point. I've thought for a while now that one of the big risks of present AI chatbots is that they can endlessly validate any arbitrary thoughts of an emotionally unstable human operator, creating a feedback loop that results in a strong irrational conviction, that might even be harmful to the operator or to others. (See: Lemoine and LaMBDA, or the person you linked to.) So I guess the risk there would be, the AI…

As a hobbyist AI-risky person, I would welcome this as an improvement on the status quo, with two worries:

- it would create a false sense of safety that we "have the problem handled"

- it would not, I believe, do much to slow down capability research or shift the ratio of capability research to safety research, as capability research is largely institutional and not dependent on public prompt access.

Ie. analogously, our main problem is not "bioterrorism", it's "gain of function research" even if the outbreak happens in public. Given an unsafe bot, not giving the public access only delays the issue. (Though delays may be valuable!)

Also, at the moment, "people interacting with GPT" is doing a lot (maybe an undue amount!) to spread awareness of AI risk. So I would welcome this as part of a comprehensive strategy, but it'd probably not be my main priority.

Though if GPT-4 has another surprise on the level of chain-of-thought or in-window reinforcement learning waiting for us, a restriction like that may be the difference between "safe" and "iffy". Tradeoffs...

Post reply on HN