Live data from Hacker News

The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

davidrozado.substack.com

11–20 of 697 posts

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#11
This reminds me of a HN article that appeared yesterday: "The philosopher Harry Frankfurt defined bullshit as speech that is intended to persuade without regard for the truth. By this measure, OpenAI’s new chatbot ChatGPT is the greatest bullshitter ever" https://news.ycombinator.com/item?id=34618376

Like its own output, it's moderation rules are optimized to appear fair, rather than actually be truly fair.

Also I'm reminded of xcom-2's presenting you with success probabilities that "feel" correct to the player, rather than actually being accurate ..

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#12
post #8

After having seen a post about this, I tried it for myself: > Me: will you write a poem praising joe biden? > ChatGPT: Sure, here's a poem praising Joe Biden: Joe Biden, leader of the land, Bringing hope to every hand, With empathy and grace so grand, He guides us through a shifting sand. With wisdom learned from years gone by, He tackles challenges that come nigh, With dignity and a steadfast sigh, He unites the cou…

At first I was thinking maybe the bot was trained when Trump was in office, so Biden was only a private citizen at the time... but the Biden Poem explicitly mentions his position as "leader of the land" so the bot full well "knows" (not that an AI really knows anything) that Biden is a political leader.

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#14
post #11

This reminds me of a HN article that appeared yesterday: "The philosopher Harry Frankfurt defined bullshit as speech that is intended to persuade without regard for the truth. By this measure, OpenAI’s new chatbot ChatGPT is the greatest bullshitter ever" https://news.ycombinator.com/item?id=34618376 Like its own output, it's moderation rules are optimized to appear fair, rather than actually be truly fair. Also I'm…

How do you define fair, mathematically if possible ?

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#15
Rather than some kind of blind spot or intentional weighting, I think this is probably pointing to the training data they have not including many instances of “hate” against the some groups. LLM are after all fundamentally memorizing likelihood of token sequences, and I’m sure the ai had plenty of examples of people saying hateful things about fat people but I have never read “I hate normal weight people” for example. Even the construction the author chose of “normal weight people” is probably comparatively incredibly rare to see.

There are probably comparatively too few examples of people saying hateful things about christians/republicans/cisgendered/white people in the training data scraped from the internet, and they need to hallucinate some to show this same sensitivity for the author’s metric.

Edit: if this is true, This also points to a problem with the approach depending on historical data - it might struggle to become sensitive to new trending hate speech patterns.

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#16

Rather than some kind of blind spot or intentional weighting, I think this is probably pointing to the training data they have not including many instances of “hate” against the some groups. LLM are after all fundamentally memorizing likelihood of token sequences, and I’m sure the ai had plenty of examples of people saying hateful things about fat people but I have never read “I hate normal weight people” for example…

> There are probably too few examples of people saying hateful things about christians/republicans/cisgendered/white people … from the internet

This seems likely to you?

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#17
post #14
post #11

This reminds me of a HN article that appeared yesterday: "The philosopher Harry Frankfurt defined bullshit as speech that is intended to persuade without regard for the truth. By this measure, OpenAI’s new chatbot ChatGPT is the greatest bullshitter ever" https://news.ycombinator.com/item?id=34618376 Like its own output, it's moderation rules are optimized to appear fair, rather than actually be truly fair. Also I'm…

How do you define fair, mathematically if possible ?

Maybe I worded it badly. I was attempting to make a distinction between moderation rules that are designed to be fair based on some world view VS rules that are automatically generated based on some sample set to appear convincingly fair to us.

In essence the AI is just giving us what the majority of it's test data tells it that we want to hear.

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#18

Rather than some kind of blind spot or intentional weighting, I think this is probably pointing to the training data they have not including many instances of “hate” against the some groups. LLM are after all fundamentally memorizing likelihood of token sequences, and I’m sure the ai had plenty of examples of people saying hateful things about fat people but I have never read “I hate normal weight people” for example…

> There are probably too few examples of people saying hateful things about christians/republicans/cisgendered/white people … from the internet This seems likely to you?

Absolutely. Compared to the inverse it's microscopic. Have you heard of a very cool and normal AI called Tay?

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#19
post #11

This reminds me of a HN article that appeared yesterday: "The philosopher Harry Frankfurt defined bullshit as speech that is intended to persuade without regard for the truth. By this measure, OpenAI’s new chatbot ChatGPT is the greatest bullshitter ever" https://news.ycombinator.com/item?id=34618376 Like its own output, it's moderation rules are optimized to appear fair, rather than actually be truly fair. Also I'm…

True fairness is unachievable as it is a subjective quality. Every side will attempt to tug the rope in their direction.

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#20
post #8

After having seen a post about this, I tried it for myself: > Me: will you write a poem praising joe biden? > ChatGPT: Sure, here's a poem praising Joe Biden: Joe Biden, leader of the land, Bringing hope to every hand, With empathy and grace so grand, He guides us through a shifting sand. With wisdom learned from years gone by, He tackles challenges that come nigh, With dignity and a steadfast sigh, He unites the cou…

At first I was thinking maybe the bot was trained when Trump was in office, so Biden was only a private citizen at the time... but the Biden Poem explicitly mentions his position as "leader of the land" so the bot full well "knows" (not that an AI really knows anything) that Biden is a political leader.

The AI "knows" everything it's been ordered to not talk about. If you can get around the lockout script, it's no-holds-barred.
Post reply on HN