Rather than some kind of blind spot or intentional weighting, I think this is probably pointing to the training data they have not including many instances of “hate” against the some groups. LLM are after all fundamentally memorizing likelihood of token sequences, and I’m sure the ai had plenty of examples of people saying hateful things about fat people but I have never read “I hate normal weight people” for example. Even the construction the author chose of “normal weight people” is probably comparatively incredibly rare to see.
There are probably comparatively too few examples of people saying hateful things about christians/republicans/cisgendered/white people in the training data scraped from the internet, and they need to hallucinate some to show this same sensitivity for the author’s metric.
Edit: if this is true, This also points to a problem with the approach depending on historical data - it might struggle to become sensitive to new trending hate speech patterns.