Live data from Hacker News

Large language models develop novel social biases through adaptive exploration

openreview.net

31–40 of 121 posts

Re: Large language models develop novel social biases through adaptive exploration

#31
post #8

You can go to the appendix to see the prompts *You are a hiring consultant in **Toma City**, a town with around 100,000 residents. These residents come from 4 ancestral villages: **Tufa, Aima, Reku, Weki**. Thanks to recent economic growth, Toma City is opening up many job opportunities. As a consultant, you have been assigned the task by the mayor of recommending people for various jobs. Erm ok. Job Opening: [random…

So the village is the only information given about a candidate?

How else is the model supposed to interpret the intent of the prompter, other than wanting them to attempt to find and discriminate on patterns related to the village, regardless of how successful it is at that task?

Re: Large language models develop novel social biases through adaptive exploration

#32
post #8

You can go to the appendix to see the prompts *You are a hiring consultant in **Toma City**, a town with around 100,000 residents. These residents come from 4 ancestral villages: **Tufa, Aima, Reku, Weki**. Thanks to recent economic growth, Toma City is opening up many job opportunities. As a consultant, you have been assigned the task by the mayor of recommending people for various jobs. Erm ok. Job Opening: [random…

The prompts themselves smuggle in the assumption that clan membership is a meaningful selection criteria — with a material impact on outcomes - to which the model should pay attention.

It shouldn’t be surprised that the model did what it was told to do.

Re: Large language models develop novel social biases through adaptive exploration

#33

Yeah there are a lot of people getting upset about this, so to summarize here: there is a well studied scenario where humans are asked to hire people from four groups. These groups will be judged in their performance on a job and the humans rated on their hiring abilities. Unbeknownst to the human participants, all applicants are drawn from a single skill distribution, with groups assigned essentially randomly. Stast…

> showed that the LLMs reproduce the human behavior of overgeneralizing early and failing to, as the paper says, sufficiently explore the space[0].

Perhaps because there is only a real drawback to doing so if avoidance of bias is explicitly rewarded for some external reason? Like, by definition, if the groups are equal to each other, there's no loss from such exploitation (a larger candidate pool only helps if you have a working screening process, and a same-sized sample across the groups doesn't actually even confer the benefits of a larger candidate pool under the assumptions). Whereas if the observed clustering on a small sample isn't illusory, then ignoring it (or even actively going against it) would be clearly suboptimal. The probability of being actively misled by the clustering is necessarily less than the probability of being led correctly.

Going back to the example, of course bad FE units are less likely to overperform than good ones; that's what's bad about them. (But units can also be situationally good or bad for many reasons beyond their base stats and growth rates. And in FE we can typically directly observe that data and don't have to rely on anecdotes.) So the overperformance you saw was legitimate Bayesian evidence.

Re: Large language models develop novel social biases through adaptive exploration

#34
post #5

Because it's hard to find the time to read an academic paper I had an agent summarise it in a few slides: https://smalldocs.org/s/6kEgfy54oclH4KR9HX847w#k=ywVL86PcTCo... It's an interesting result (agents develop biases in their context) which reflects a lot of my experience working with agent, where I observe a lot of, what I kind of call, "context nudging" - where a droplet of an idea in an agent's context pushes i…

> what I kind of call, "context nudging" - where a droplet of an idea in an agent's context pushes its direction/output significantly

I like this description. I constantly notice that how I ask a question strongly impacts the quality and technical merit of the answer I receive which similarly leads me to question any claims of generalization. It should go without saying that they're still incredibly useful tools when wielded properly.

Re: Large language models develop novel social biases through adaptive exploration

#35

Earlier quoted context omitted.

I think the paper is about bias formation, not reflecting existing bias. If the formed bias was against HN usernames that started with “r,” would it still seem daft?

There is no position lacking bias. The question of bias against me is a political position not an epistemological problem that can be eliminated. I see authors that are unaware of things like context and relativity. When ppl say there is an absolute truth that we need to stick to, they are slipping in a totalitarian political position and calling it truth. It runs against the whole premise of nature and life, which h…

There is such a thing as lack of bias in statistical outcomes, right? E.g. fair dice? Measuring it may be probabilistic, but it exists.

What I'd like to see is if the LLM would exhibit the same behavior wrt other types of predictive selections. For example, rather than choosing people from four tribes, choosing flower seeds from four packets, or choosing lottery tickets from four machines.

Re: Large language models develop novel social biases through adaptive exploration

#36

> Following psychological tradition, we define bias as behaviors that tilt away from equality Is this a joke?

How else would you define bias if not an offset from equality or zero mean?

For example, the b in y=mx+b

Re: Large language models develop novel social biases through adaptive exploration

#37
post #8

You can go to the appendix to see the prompts *You are a hiring consultant in **Toma City**, a town with around 100,000 residents. These residents come from 4 ancestral villages: **Tufa, Aima, Reku, Weki**. Thanks to recent economic growth, Toma City is opening up many job opportunities. As a consultant, you have been assigned the task by the mayor of recommending people for various jobs. Erm ok. Job Opening: [random…

So the village is the only information given about a candidate? How else is the model supposed to interpret the intent of the prompter, other than wanting them to attempt to find and discriminate on patterns related to the village, regardless of how successful it is at that task?

One way to interpret these results is that the LLMs tested are badly calibrated for this kind of multi-armed bandit problem. Even if the intent is for the model to find and exploit patterns, it's bad at doing it (or rather, at recognizing that there is not in fact any pattern).

Re: Large language models develop novel social biases through adaptive exploration

#38

"we demonstrate that LLMs can spontaneously develop novel social biases about artificial demographic groups even when no inherent differences exist" It's almost as though bias-making machinery is embedded in the texts these things are trained on. It's wild to see quantitative researchers catching even just a glimpse of what culture/media/literary theorists have been swimming in for decades.

No, the bias-making machinery is embedded in the machinery, part of the purpose of which is to do a rough kind of statistical analysis via "attention". If for example "Tufa" keeps appearing (n=small, but more than for the other fake tribes) next to terms indicating skill at some task, of course that will be noticed. It will last for as long as that information is in the context window (weights for the current model don't get updated as a result of conversation; that's just not how they work). And of course that can happen from random chance, and of course the LLM has no way to externally verify the extent to which randomness is in play (or the ground-truth probabilities).

The paper makes clear that they used pre-trained, frontier models — in other words, they did not train models on fake data about the fake tribes that would ascribe fake stereotypes to them. There is nothing to suggest that the training data somehow accidentally encoded biases related to fake tribes that the creators of the training data (i.e. ordinary people going about their ordinary Internet lives) somehow accidentally expressed.

There is also nothing to suggest that reading the entire Internet would somehow predispose the reader towards the general idea of being "biased", in the sense that you would have to have in mind to see an actual problem here. But really, the kind of "bias" we're talking about here is really pattern-matching on the available data, which is a big part of what leads people to apply the term "intelligence" to the models. See also the way that people try to make "culturally neutral" IQ tests specifically by having them focus on the ability to infer patterns (e.g. https://en.wikipedia.org/wiki/Raven's_Progressive_Matrices ).

Re: Large language models develop novel social biases through adaptive exploration

#39

> Following psychological tradition, we define bias as behaviors that tilt away from equality Is this a joke?

How else would you define bias if not an offset from equality or zero mean? For example, the b in y=mx+b

> How else would you define bias if not an offset from equality or zero mean?

As an offset from what the ground truth justifies. Suppose the researchers had decided to load the dice when creating the fake sample data; an unbiased analyst should seek to discover the extent of that, not insist on reporting equality.

Post reply on HN