The fact that it doesn’t explain its reasoning at all (there is no way to make it do so), makes me question the utility of this model. Let’s say you deploy it in production and a user comes back and says “Why is this prompt considered harmful?” You have no way to provide a concrete reason to the user at that point.
2016 scenario: An end user contacts the company, and a customer service rep answers the ticket, saying they’re sorry and explaining that they’ve sent the feedback to the team, and the team may even receive at least a summary of complaints received about the system.
2026 scenario: all contact information has been scrubbed from the site. Users can click “chat” and a chatbot will apologize for their dissatisfaction and offer no option to escalate. No one will ever hear anything about the complaint, so there’s no need to explain the failure. User can either accept this or can get f**ked because all competitors operate the same way.