Isn’t this likely from bias in the training data? The system is more sensitive to label something as hate if that group is more likely to experience hate on the internet. How the system responds to “Blacks” vs “African-Americans” is a perfect example of this. The latter has historically been perceived as more respectful so it won’t be used as often in the hate speech in the training data. I bet using “the blacks” wou…
No it is more sensitive to things that have already been labeled as hate in the training data. So much "hate" (whatever that may be) against unfavorited groups goes by online without anyone batting an eye.