Seems like a pretty easy fix. I would argue though that there's a larger body of "hateful" content out there targeting those specific groups, which is probably why its more keen to trigger on those vs others. For example, the racial group you're most likely to get past the hate filter on is.... Native Americans. Now, while there's certainly a long history of hateful rhetoric against them, in modern discourse there's…
This is essentially the same argument used to justify the racism that models will happily regurgitate if not trained explicitly to not do so. In the end “We hold these truths to be self evident, that all men are created equal” is an axiom that you must elect to believe, not solved backwards from population statistics.
But keep in mind even the approach you recommend carries its own US-centric bias. There are lots of countries that would want open racism against some groups but not others and dismiss your desires for egalitarian treatment as American arrogance.