Earlier quoted context omitted.
If someone gives a language model the capability for unfettered interaction with the physical world, and they are not liable for the consequences, then no safety feature of Claude can save us. And if they are liable, that is the primary mechanism which will ensure they take necessary steps to avoid negative consequences.
How many years until we have an AI decision making system that requires human approval convincing the human its decision is best, ending in catastrophe? (Sorry, posed merely for thought and not for dismissal of current/future achievements.) Makes me also wonder if you had two polarized bots arguing/discussing with each other, how long would it take for one to convince the other?
However, the AI is but one of many agents trying to convince me, and there are many other things also in the mix. It is a bit like banning the knowledge of the second world war in fear that someone will learn that fascism is possible.
For this whole song and dance we call civilization as we know it to function, we have to say humans are accountable for their actions, with some well defined exceptions like duress and insanity. Save those exceptions, I can't absolve my liability by saying something, or someone, convinced me to do something by talking to me.
If ChatGPT comes to me with a gun to my head and tells me to do something, that is a different matter, but then liability shifts to whoever gave ChatGPT a gun, even if ChatGPT convinced that person to give it a gun by making a really good written argument.