Marc Andereseen has talked about the downside of RLHF: it's a specific group of liberal low income people in California who did the rating, so AI has been leaning their culture. I think OpenAI tried to diversify at least the location of the raters somewhat, but it's hard to diversify on every level.
What do low income people have to do with it, when AI companies and research is borne out of Silicon Valley culture of rich, liberal Californians? I'm still waiting for models based on the curt and abrasive stereotype of Eastern European programmers, as contrast to the sickeningly cheerful AIs we have today that couldn't sound more West Coast if they tried.
RLHF is "ask a human to score lots of LLM answers". So the claim is that the AI companies are hiring cheap (~poor) people from convenient locations (CA, since that's where the rest of the company is).