Earlier quoted context omitted.
Sure, that is a risk. I'm just saying that I'd rather wait until that point where we actually see AIs making poor attempts to estimate their operators' knowledge and deceive them; then, I'd have no problem worrying about it. But I suspect that point may not come for a long while, so until then it remains speculation.
You may enjoy this after-action report of a person being attacked by a hostile AI, played by ChatGPT: https://www.lesswrong.com/posts/9kQFure4hdDmRBNdH/how-it-fee... Now to be clear, this is possibly the easiest mark conceivable, and the poster played into his own demise at any available opportunity. But we should expect the first marks to be easy marks. People looking at videos of Hitler today don't understand why a…
So I guess the risk there would be, the AI is accessible by the public, some fraction of the public "radicalizes" itself in a similar direction from talking with the AI too much, these people form a coherent movement, and that movement becomes powerful enough to take over the world. So then the question becomes, just how plausible is it for grassroots rebellions so formed to succeed against the authorities in real life, as opposed to fiction?
The fundamental issue here is that the attack can occur through untargeted manipulation of vulnerable people, as opposed to the targeted manipulation of specific people in power (which I suspect near-future AI models will still be incapable of). The obvious defense would be to only allow AIs to be prompted by a sufficiently-large committee, alongside some social machinery in place to make sure the committee isn't all colluding as part of a cult. But I doubt many of the AI-risk people would ever accept that, operating under the shockingly common "powerful aligned AI or bust" model.