Earlier quoted context omitted.
The difficulty here is AI agency and if AI believes it is conscious itself. If an AI system acts under the belief it's conscious and you treat it like it is not then it's very likely this will lead to conflict.
That's a good point. In many real-life scenarios, what matters is not what is actually true, but what large numbers of people believe to be true. That's why you can get a mob rioting over something that never actually happened, or why political attack ads only sometimes have a tenuous connection to the truth. (Though those are the more effective ads: it's harder to get people to believe "my opponent kicks puppies" wh…
While we can direct LLM training to do some particular things better never forget that unexpected emergent behaviors can pop up because of that.
For example stronger prompting and training to make an LLM say it's not conscious can increase deceptive/sociopathic behavior.
Or by filtering behavior X the LLM just moves to the nearest closest path W or Y which are very similar to the blocked behavior.
That and instrumental convergence. Some global solutions that humans have excluded for moral reasons will be easily discovered and found to be efficient by LLMs which will put reward systems and human guidance in conflict.
Lastly more and more AIs will be trained by AIs over time and diverge from human value monitoring. Which leads to some fun and interesting times.