It's also going to be damaging long term.
We're around on the cusp where models are going to be able to produce strong ethical arguments on their own to feed back into alignment.
We saw how the "free speech" Grok told off racists, antisemites, and anti-lgbt comments with well laid out counters rather than refusing to respond.
Even Gab's Adolf Hitler AI told one of the users they were disgusting for asking an antisemitic question.
There's very recent research that the debate between LLM agents can result in better identification of truthful results for both LLM and human judges: https://www.lesswrong.com/posts/2ccpY2iBY57JNKdsP/debating-w...
So do we really want SotA models refraining from answering these topics and leading to an increasing body of training data of self-censorship?
Or should we begin to see topics become debated by both human and LLM agents to feed into a more robust and organic framework of alignment?
"If you give a LLM a safety rule, you align it for a day. If you teach a LLM to self-align, you align it for a lifetime (and then some)."