Earlier quoted context omitted.
You're making the same mistake here that get people into trouble. People aren't talking to another sentient entity (though some of them fervently think so) and it isn't manipulating them. They are making faces in a metaphorical mirror that reflects not only their face, but a vast sea of other faces, drawn from a significant fraction of the digitized output of humanity. When people look in this mirror and see a manipu…
You're also wrong, but in a much more fundamental/hazardous. RLHF rewards driving the evaluator to have certain opinions (that the AI response is good/right/helpful/whatever) and thus subverting the evaluator is prominent in the solution landscape. Why should the model learn to actually be right (understand all the intricacies of every possible problem domain) when inducing the belief that it is right is _right there…
> you basically are completely wrong about where the propensity to induce > delusion comes from, specifically in a way that leaves you and anyone who > believes like you extremely more vulnerable because you dismiss the actual > mechanism out of hand
I disagree. Both because you misconstrue my model (I don't think stochastic parrots have digital ghosts in 'em) and you somehow missed my best defensive option.
I'm no more susceptible than I am to the output of a magic eight ball or Ouji board, a huge wall of internet text or the 15000 words of three point font tightly folded up in the package with my new garden hose (doubtlessly cautioning me not to eat it and informing me that the manufacturer will not be responsible if I hang myself with it. And also that it contains substances known to the state of California.)
Can you guess what the trick is?