I don't think AI needs to be able to explain itself in order to be trusted. Human beings, in general, cannot give arbitrarily deep and valid explanations for their actions, and yet, somehow, we manage to come to trust them.
How do you apply the same logic to a neural net? It has no comprehension of desirable outcomes, or what its choices actually entail and never can.
There is no way to tell a neural net that a human is not a signpost directly, so it knows in the future no humans are signposts. The only thing you can do is train it on a larger data set and hope it "gets it".
You can tell a human being that a human is not a signpost once, and they can abstract it to all situations.