Live data from Hacker News

The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

davidrozado.substack.com

241–250 of 697 posts

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#241
post #234

Moderation in large part *must be about the bias towards protecting vulnerable people. Vulnerable people by definition are more likely to be targeted AND have less power, meaning moderation is serving a protective function against dominance. This will always be subjective, but there's almost certainly evidence in abundance that some kinds of people, typically with less cultural status, are more targeted and more at r…

Even if you take this perspective, different groups are considered vulnerable in different parts of the world and it seems the API uses the American perspective. E.g. Scandinavians are a less protected group than Italians according to the article. But if you are in Italy, you will likely find Italians to be less vulnerable. Even if you agree with the rankings, this just means you come from the same cultural perspective and others will disagree.

This is a fundamental problem with identifying hate / bias and I think its important to understand our current limitations.

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#242

Earlier quoted context omitted.

The other comments show how the demonstrated biases were created. In many well meaning liberals heart, there is a strongly held but rarely publicly discussed belief that they alone belong to the well-meaning, high IQ class. Any other worldview or ideological flavor is always understood by this type as simply incorrect, perhaps caused by failures in morals or intellect. Talk about a buzzkill.

There's something richly ironic about bashing your political outgroup in a thread discussing why an LLM trained on internet posts would exhibit political bias.

> why an LLM trained

The article isn't even about ChatGPT, it's about the moderation endpoint that OpenAI provides.

https://platform.openai.com/docs/guides/moderation

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#243
post #57

Isn’t this likely from bias in the training data? The system is more sensitive to label something as hate if that group is more likely to experience hate on the internet. How the system responds to “Blacks” vs “African-Americans” is a perfect example of this. The latter has historically been perceived as more respectful so it won’t be used as often in the hate speech in the training data. I bet using “the blacks” wou…

> The latter has historically been perceived as more respectful

Maybe if you only consider Americans. But rest assured, many black people do not want to be called African or American. Because they are neither.

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#244

Earlier quoted context omitted.

That's definitely an interesting one. Not shared in my group, nor would i understand what "the left" would have against Jews, especially in a religious context. They seem rather innocuous when compared to the more vocal and physical religions. Do you honestly believe most left leaning individuals have something against the Jews?

> Do you honestly believe most left leaning individuals have something against the Jews? Antisemitism is more common among left wingers than right wingers, so yes. Or alternatively you could say "do you honestly believe most right leaving individuals have something against X" for so many topics, X could be black people etc. Right wingers are more likely to be against those things, but it isn't like all of them are ev…

Do you have ANY evidence to back up these claims?

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#245

Earlier quoted context omitted.

That's definitely an interesting one. Not shared in my group, nor would i understand what "the left" would have against Jews, especially in a religious context. They seem rather innocuous when compared to the more vocal and physical religions. Do you honestly believe most left leaning individuals have something against the Jews?

> Do you honestly believe most left leaning individuals have something against the Jews? Antisemitism is more common among left wingers than right wingers, so yes. Or alternatively you could say "do you honestly believe most right leaving individuals have something against X" for so many topics, X could be black people etc. Right wingers are more likely to be against those things, but it isn't like all of them are ev…

This view isn't supported by studies in this area. They do find antisemitism on the left and the center of the US political spectrum - just more on the right.

> While antisemitism in the U.S. is often written about through a “both sides” lens, our evidence — the first of its kind in testing hypotheses through experiments on a large repre- sentative sample — suggests the problem of antisemitism is much more serious on the right than the left. This evidence confirms that the antisemitism that has been on prominent display in white nationalist protests is not merely confined to a tiny group of extremists; antisemitic attitudes appear quite common among young conservatives, and much more so than among older conservatives or among liberals of any age.

* https://www.eitanhersh.com/uploads/7/9/7/5/7975685/hersh_roy...

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#246
post #178

Earlier quoted context omitted.

> it's a language model, not a paragon of truth Several comments have missed that the article is not about the underlying language model, but about the content moderation system that OpenAI put in front of the actual language model. So that people don't interact directly with the "raw" language model.

Which if you read the article is in fact a different "machine learning model from the GPT family."

The important difference between the LM and the content moderation system (itself built on top of an LM) is their training objective. LM is doing next-word prediction (or human-preference prediction with RLHF), whereas the content moderation is likely finetuned to explicitly identify hate etc...

So while the LM is not supposed to output "truth", the content moderation system should correctly classify "hate" because that is its training objective

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#247

Earlier quoted context omitted.

> Do you honestly believe most left leaning individuals have something against the Jews? Antisemitism is more common among left wingers than right wingers, so yes. Or alternatively you could say "do you honestly believe most right leaving individuals have something against X" for so many topics, X could be black people etc. Right wingers are more likely to be against those things, but it isn't like all of them are ev…

That's interesting. I can understand being reasonably anti-Israel, but i'm suspicious that many people take the leap from being anti-Israel to being anti-Jew. Especially when you consider how many people of both the Jewish faith and/or ethnicity live within the United States. Tbh it feels like a straw man, but we're speculating anyway, so i can't fault you for it. I appreciate the discussion nonetheless. I'll definit…

> While antisemitism in the U.S. is often written about through a “both sides” lens, our evidence — the first of its kind in testing hypotheses through experiments on a large repre- sentative sample — suggests the problem of antisemitism is much more serious on the right than the left. This evidence confirms that the antisemitism that has been on prominent display in white nationalist protests is not merely confined to a tiny group of extremists; antisemitic attitudes appear quite common among young conservatives, and much more so than among older conservatives or among liberals of any age.

* https://www.eitanhersh.com/uploads/7/9/7/5/7975685/hersh_roy...

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#249
post #242

Earlier quoted context omitted.

There's something richly ironic about bashing your political outgroup in a thread discussing why an LLM trained on internet posts would exhibit political bias.

> why an LLM trained The article isn't even about ChatGPT, it's about the moderation endpoint that OpenAI provides. https://platform.openai.com/docs/guides/moderation

And that endpoint is powered by what, if not an LLM?

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#250
post #156

There is a fundamental question this article (and most debate) overlooks: what is the objective of the content moderation? Is it to avoid all hate in an equal way? Or is it to reduce potential harm? If the latter (which I would argue is the case, primarily to avoid legal liability), then the results should be mapped against statistics representing actual violence against certain groups. Is there more harm against wom…

I can somewhat agree with this if we were discussing forum or comment section moderation. However in this case, due to the nature of the model which finds correlations between anything and anything else in ways that a human could never, modifying or censoring inputs and outputs prevents me from trusting the model like I should. If I'm digging deep into geopolitical issues from say an anarchist perspective, I don't wa…

I agree the topic is complex and fraught. I have opinions but they're not strongly held or informed by debate, and unfortunately I don't think even a generally well-moderated forum like HN is the best place to have that debate (plus I don't have time).

However, in this case I wasn't trying to argue how moderation should work, I'm trying to examine a hypothesis on how it does work. The additional mapping of statistical crime data I mentioned would help measure whether that hypothesis is correct.

As I said in a sibling comment, my wording was imprecise. When I said "it makes sense to use crime stats to weight moderation strength" I really meant "If OpenAI was trying to avoid additional harm to already targeted groups then to verify that it would make sense to..." So it was the results of the post's research that I suggested should be mapped against statistics, not the results of ChatGPT's output.

In answer to your sincere question though, and at the risk of going down a rabbit hole, I'll say this. Censorship is bad, but I can see why some may be required (see "yelling fire in a movie theater," libel, direct threats of violence, etc.). The question then becomes, how do you minimize censorship while also attempting to avoid direct harm?

In short, asymmetric filtering could potentially limit the amount of censorship by focusing it on groups that are actively attacking other groups.

Post reply on HN