There is a fundamental question this article (and most debate) overlooks: what is the objective of the content moderation? Is it to avoid all hate in an equal way? Or is it to reduce potential harm? If the latter (which I would argue is the case, primarily to avoid legal liability), then the results should be mapped against statistics representing actual violence against certain groups. Is there more harm against wom…
I can somewhat agree with this if we were discussing forum or comment section moderation. However in this case, due to the nature of the model which finds correlations between anything and anything else in ways that a human could never, modifying or censoring inputs and outputs prevents me from trusting the model like I should. If I'm digging deep into geopolitical issues from say an anarchist perspective, I don't wa…
> Why can we not consider all groups equal?
Because they aren't. (If that's a thing you find debatable, LMK!) Attempting to consider them all equal runs into problems just like you'd run into problems trying to consider all pumps at the gas station (including diesel) equal; even if, for sake of the metaphor, all the prices were equal.
> why should the results be mapped against statistics
I'm not sure that's answerable outside of a specific context. Personally, I think we should because it's fascinating and, I would expect, a really informative way to explore in more detail. Culturally, it's because you get better communities when you do things like give the high-accessibility seat on the train to the person with broken leg; aka, go out of your way to treat harmed people with more care.
If your question is about the inverse, something like "why should language that's harmful against one group by OK when used against a group that doesn't experience it harmfully", I dunno what to say. I feel like the question kinda answers itself; sort of a "ain't broke don't fix".
Or maybe your question is more: "why does this disagree with [me] about what's harmful to [me]", I dunno. Maybe the training data didn't include (enough) for the AI-as-sensor to detect that, maybe that data doesn't exist, maybe (and very cynically) that data doesn't actually point to that "conclusion" when run through the AI. Kinda like how the first time many people experience delayed-onset-muscle-soreness they think it's "I'm hurt"-pain.