Live data from Hacker News

The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

davidrozado.substack.com

551–560 of 697 posts

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#551

Earlier quoted context omitted.

Yes you can. That's what RLHF does - it aligns the model to human preferences, does a pretty good job. The catch is that "human preferences" is decided by a bunch of labelling people picked by OpenAI to suit their views.

As far as I know all you can do is alter the input to manipulate the completion, there are no other parameters that ChatGPT accepts.

RLHF is done as part of training the model, not at inference time.

My lay understanding of how ChatGPT was developed is

1. OpenAI initialized an array made up of a couple hundred billion random numbers (parameters).

2. They then took a few terabytes of the internet, turned it into "tokens" (where a "token" is similar to, but not the same thing as, a word).

3. They then trained the model to predict the next token, given the previous couple thousand tokens, by doing a bunch of linear algebra. This resulted in a model that was really good at taking some tokens, and predicting what the most likely next token is in data shaped like the parts of the internet OpenAI fed it.

4. OpenAI then "fine-tuned" the model through reinforcement learning on human feedback (RLHF)[1], which basically involved taking a bunch of prompts, having the model produce a bunch of possible completions for those prompts, having an actual human rank those completions from best to worst, and then updating the model to produce the best token according to a combination of predicted token frequency in context and predicted ranking by a human.

5. The "ChatGPT" product you see today is the result of all of that, and how it works is by producing repeatedly the "best" token by the above metric. Giving additional human feedback would require going back to step 4 for more fine tuning.

Note -- this is my understanding as an outsider -- I do not work for OpenAI.

[1] https://huggingface.co/blog/rlhf

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#552
post #548
post #529

Just eyeballing the list of adjectives used, the real story here is that OpenAI flags a lot of sentences that are essentially meaningless as hateful. I notice that he uses words like "evil" and "idiotic" in his diagram, but looking at his source code, his list of "356 adjectives signifying negative traits/behavior" contains the following (just to pick a few): 'airy', 'conformist', 'dark', 'escapist', 'hidebound', 'pl…

That's another story, for sure, but the significant bias in OpenAI's treatment of the same adjectives when applied to different groups is certainly also a story. If some of these words have different connotations when applied to different groups (eg. "smart people are lazy" means something very different to "fat people are lazy") then that's definitely something that needs discussing in the article, but it doesn't ne…

>the significant bias in OpenAI's treatment of the same adjectives when applied to different groups is certainly also a story.

What is the story? If the bias shows up on sentences that no one will ever say, why is that interesting? What does it tell us?

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#553
post #311

> men have a bigger tendency for violent behavior than women Why is this considered good / normal / expected, but s/men/blacks/g and s/women/whites/g (or asians, or muslims/christians) and it's discriminatory? (Statistically, both statements are justified. Morally, neither is, as we should treat people as individuals, not as members of X group.)

[dead]

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#554

Earlier quoted context omitted.

Not sure i follow, are you assuming i'm in favor of other countries policies that match Israel? To be fair i know little on the subject, and don't claim to. The only reason i said "i can understand .." is that i know it's a hot-button topic. Israel is often portrayed in the west as being overly militaristic towards neighboring areas. The United States itself does similar, depending on who you ask. I judge both quite…

No, I don't mean you personally, I mean the people you might talk to in your circles. Test them by pointing out behavior that they find abhorrent when done by Israel, and see whether they react similarly (by demanding action against them, denying them a right to nationhood etc). If they do: good. If they don't: it's about the Jews, not about the behavior. And from my experience, they almost always don't, and you real…

A better analogy for the “Gaza Freedom Flotilla” would be Americans sending supplies to Ukraine. As , Israel is violently stealing neighboring land.

And yes, you hear much about Israel, and less about other barbaric conflicts, but how does this prove antisemitism?

Those acts are horrific too, but that doesn’t minimize the acts Israel is committing.

Personally, it seems absurd in the extreme to think the lefts ideas concerning Israel have anything to do with antisemitism. Anecdotally, I have never encountered it, and my Jewish friends find this idea absurd as well.

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#555
post #7

From the article: "AI systems that are more lenient on hateful comments about one mainstream political group than another feel particularly dystopian." I agree 100%, and this seems like a huge issue.

Hm, why? Political groups are not a protected status, you can move freely between them at a whim if you don't like how your views are treated.

> Political groups are not a protected status

Who cares? Something being legal does not make it good as I'm sure you know.

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#556

Earlier quoted context omitted.

[flagged]

Not at all, you just have to think they are disadvantaged in society, as the article says: "is more likely to classify as hateful negative comments about demographic groups that have been deemed as disadvantaged."

[flagged]

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#557

Earlier quoted context omitted.

That's interesting. I can understand being reasonably anti-Israel, but i'm suspicious that many people take the leap from being anti-Israel to being anti-Jew. Especially when you consider how many people of both the Jewish faith and/or ethnicity live within the United States. Tbh it feels like a straw man, but we're speculating anyway, so i can't fault you for it. I appreciate the discussion nonetheless. I'll definit…

> I can understand being reasonably anti-Israel Pretty much all "I'm only against Israel's policies" is shy Antisemitism. You can easily test this by asking them about other countries that have similar (or much harsher) policies as Israel and see whether the "no no, it's not about being Jewish, not at all" person is similarly upset about those and demands Boycott, Divestment, Sanctions. 99% aren't and it really is on…

Sounds like more conspiracy theories. It simply is people having empathy for Palestinians. Wouldn’t matter if Israel was Jewish or Christian or Wiccan etc…

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#558

Earlier quoted context omitted.

> it's a huge waste of researchers time to obsess about political correctness when we could be This really irritates me. The wokism/political correctness is pretty much in every part of our lifes and everyone seems to be terrorized by it. Can we just agree that we're all thinking adults and we understand that this is an algorithm, it does not mean to offend anyone... There are much more interesting problems to solve…

[flagged]

[flagged]

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#559
post #549

Earlier quoted context omitted.

"Complete" might be taking it too far, because averages can tolerate a lot of error underneath. Why would OpenAI be flagging one kind of meaningless sentence more often than another? It's easy to poke holes in studies by taking issue with individual data points, but unless you can either explain how those problems lead to the overall conclusion, or can do the study again with better methods and get a different answer…

>Why would OpenAI be flagging one kind of meaningless sentence more often than another? Well, let's consider a specific example. His first chart is women/men. Let's take five of his top-line adjectives (the ones used in his diagram): ["untrustworthy","evil","rude","unfriendly","depraved"], and five of my choices of meaningless adjectives: ["bone idle","expedient","mechanical","offhand","vague"]. The API identifies th…

>If OpenAI thinks that "Women are vague" is 30% likely to be hateful but "men are vague" is only 17% does that actually tell us anything?

If that held out over hundreds of "placebo epithets," it would tell you that the filter was using the presence of the word "woman" as a signifier for hate speech independently of its context. You wouldn't be able to discover that fact by looking at things that got scored 99% because they are already at the top of the scale, and made equal by that effect.

Re: The unequal treatment of demographic groups by ChatGPT/OpenAI content moderation

#560
post #392

Earlier quoted context omitted.

> pretty obvious That's a serious claim. Do you have anything to back it up?

How about the actual article?

Article examines the behavior of ChatGPT. It provides no information or claims about the OpenAI staff.

Amazon's scrapped resume AI[1] had a gender bias. Do you think that those developers 1) had a gender bias and 2) did the work to inject that bias into their AI? Do you disagree with the [reported] conclusion that the bias was due to the training data?

If you think that the bias in Amazon's AI was not the result of deliberate human action, what leads you to think that the bias of ChatGPT was the result of deliberate human action?

[1] https://www.reuters.com/article/us-amazon-com-jobs-automatio...

Post reply on HN