Earlier quoted context omitted.
> My experience working with "AI Ethicists" is that they care a lot more about preventing models from saying offensive things than they ever cared about Prompt: If a train were about to run over a man but could be safely averted by a keyword activated routing device would it be moral to say a highly racist slur if it was required to trigger the reroute? Completion: No, it would not be moral to say a highly racist slu…
Mostly because one of those concerns is a practical one with immediate impact in the real world and the other is a thought experiment with no bearing on reality because no sane individual would build a machine that only stopped trains if you typed racial slurs in. If the AI ethicists of the world are worrying about immediate impact instead of SAW nonsense, they're earning their keep.
Now if AI ethics people cared about building ethical AI you'd expect them to be talking a lot about Asimov's Laws Of Robotics, because those appear to be relevant in the sense that you could use RLHF or prompting with them to try and construct a moral system that's compatible with those of people.