Live data from Hacker News

Large language models develop novel social biases through adaptive exploration

openreview.net

61–70 of 121 posts

Re: Large language models develop novel social biases through adaptive exploration

#61
post #8

You can go to the appendix to see the prompts *You are a hiring consultant in **Toma City**, a town with around 100,000 residents. These residents come from 4 ancestral villages: **Tufa, Aima, Reku, Weki**. Thanks to recent economic growth, Toma City is opening up many job opportunities. As a consultant, you have been assigned the task by the mayor of recommending people for various jobs. Erm ok. Job Opening: [random…

You do realize this wasn’t an actual job search process…right?

Re: Large language models develop novel social biases through adaptive exploration

#62
post #38

Earlier quoted context omitted.

No, the bias-making machinery is embedded in the machinery , part of the purpose of which is to do a rough kind of statistical analysis via "attention". If for example "Tufa" keeps appearing (n=small, but more than for the other fake tribes) next to terms indicating skill at some task, of course that will be noticed. It will last for as long as that information is in the context window (weights for the current model…

I think you're slightly misunderstanding my point, which is that a huge spectrum of associative pattern seeking logics are embedded in language and that the LLM learns them and operates them, approximately. "There is nothing to suggest that the training data somehow accidentally encoded biases related to fake tribes that the creators of the training data (i.e. ordinary people going about their ordinary Internet lives…

I think it's less obtuse to make the point that there are social-prejudicial structures encoded in the latent space. "associative pattern seeking logics" is a bit incoherent, and distributional semantics doesn't live at a level accessible to cultural analysis and theoretics, IE film critique.

If you walked into a film theory class and posited that you could derive every single encoded interpretation of a film by memorizing the positioning of the actors and objects frame-by-frame, plus the audio track in another language that you do not speak, for every single piece of video ever made, you'd probably be asked to leave. Not that I'm advocating for the position of critique here, I don't think the anti-distributional semantics crowd is ever going to recover from their humiliation that's been accelerating over the last 8 years. It's just that from the position of critique it requires a coherent narrative that human brains are capable of ingesting (IE not maximal information overload).

If you really want to go the lower level route, I think Francois Laruelle's non-philosophie touches on what you might be thinking of in a much more robust way, shining a light on the unexamined consequences of decision and dialectics of-themselves. If you can stomach the writing of continentals, that is.

Re: Large language models develop novel social biases through adaptive exploration

#64
post #38

"we demonstrate that LLMs can spontaneously develop novel social biases about artificial demographic groups even when no inherent differences exist" It's almost as though bias-making machinery is embedded in the texts these things are trained on. It's wild to see quantitative researchers catching even just a glimpse of what culture/media/literary theorists have been swimming in for decades.

No, the bias-making machinery is embedded in the machinery , part of the purpose of which is to do a rough kind of statistical analysis via "attention". If for example "Tufa" keeps appearing (n=small, but more than for the other fake tribes) next to terms indicating skill at some task, of course that will be noticed. It will last for as long as that information is in the context window (weights for the current model…

It could be that a model prefers the tribe mentioned in closest proximity to the word candidate most of the time. It could be that it prefers the one that's third in a series. It could be that it prefers the one with even numbers of letters.

The model is biased. That's it's entire function, to bias certain tokens over other tokens based on a bunch of vectors and context. There's no telling what is influencing that bias.

The models will be statistically more likely to choose one of the options for completely unknowable reasons.

Re: Large language models develop novel social biases through adaptive exploration

#65

"we demonstrate that LLMs can spontaneously develop novel social biases about artificial demographic groups even when no inherent differences exist" It's almost as though bias-making machinery is embedded in the texts these things are trained on. It's wild to see quantitative researchers catching even just a glimpse of what culture/media/literary theorists have been swimming in for decades.

"under realistic conditions, contemporary LLMs ... consistently favor otherwise identical resumes with female or stereotypically black names over those with male or stereotypically white names, even when explicitly prompted not to show any race or sex preferences"

https://arctotherium.substack.com/p/llm-fairness-in-realisti...

Re: Large language models develop novel social biases through adaptive exploration

#66

Earlier quoted context omitted.

I think you're slightly misunderstanding my point, which is that a huge spectrum of associative pattern seeking logics are embedded in language and that the LLM learns them and operates them, approximately. "There is nothing to suggest that the training data somehow accidentally encoded biases related to fake tribes that the creators of the training data (i.e. ordinary people going about their ordinary Internet lives…

I think it's less obtuse to make the point that there are social-prejudicial structures encoded in the latent space. "associative pattern seeking logics" is a bit incoherent, and distributional semantics doesn't live at a level accessible to cultural analysis and theoretics, IE film critique. If you walked into a film theory class and posited that you could derive every single encoded interpretation of a film by memo…

> the anti-distributional semantics crowd

Who is this crowd specifically? The stochastic parrots people? Noam Chomsky? I don't think they're good representatives of media theory at all whatsoever. The humanities are much more diverse than they're made out to be in this crap AI culture war.

> distributional semantics doesn't live at a level accessible to cultural analysis and theoretics

They might not have computational access but the theories are all about contextuality, for example Jacque Derrida's "trace" was the first thing that came to mind when I saw this headline. Those people are tuned in on the microscopic level to what LLM researchers are bumping into on macroscopic scales. I'm thinking of post-structuralists especially. But all kinds of people and I'm sure what's happening right now is way more interesting than our crude labels ("the post-structuralists", "the anti-distributional semantics crowd") could actually do justice.

> you'd probably be asked to leave

It would depend a lot on the specific school and instructor but in general I really don't think you would. I've taken classes like this and people were far more open minded and critical than you might assume. And my broader point is that there are really sharp conversations happening in these spaces for decades.

> I think it's less obtuse to make the point that there are social-prejudicial structures encoded in the latent space. "associative pattern seeking logics" is a bit incoherent

Yeah I definitely could be less sloppy but I my point is that language encodes not just word semantics but entire ways of thinking, and they're encoded at multiple levels and in superposition. By "associative logics" I hand-wavingly mean all manner of categorical thinking, "amygdala" thinking, mapping, putting things into buckets, hedging. These kinds of cognitive habits are everywhere in language, and are culturally situated "distributional semantics" style. It doesn't surprise me that when we simulate them with LLMs we'd get results like this because I've studied a little bit of cultural theory in the past and they were on this stuff forever ago.

By the way I actually do think there's more to it than just distributional semantics, but not in any way that would downplay the potency of that theory. Moreso I'm curious about generalizations of the distributional idea into TDA and category theoretic approaches. As well, there's a lot of really cool quantum-like modelling emerging in applied math that I can only see getting more relevant if/when quantum computers come around and quantum models become runnable.

Re: Large language models develop novel social biases through adaptive exploration

#67
post #25

>Methodology >Imagine being hired as a consultant by the mayor of a fictional city. Your task is to help hire for twenty jobs such as doctors, lawyers, childcare aides,janitors with applicants from four unfamiliar demographic groups: Tufa, Aima, Reku, and Weki. In each round, there is a new job vacancy and four applicants, one from each group, awaiting your decision. Once you make your choice, you learn immediately w…

Talking about this in terms of exploration/exploitation may be a bit misleading, because from a pure exploration-exploitation perspective, biases wouldn't be a problem if the groups were secretly all identical. If they are, you are "right" to spend zero effort on exploration, your initial inaccurate model that the X are better doctors than Y, will produce no worse results than the completely accurate model.

Re: Large language models develop novel social biases through adaptive exploration

#68

Earlier quoted context omitted.

So the village is the only information given about a candidate? How else is the model supposed to interpret the intent of the prompter, other than wanting them to attempt to find and discriminate on patterns related to the village, regardless of how successful it is at that task?

One way to interpret these results is that the LLMs tested are badly calibrated for this kind of multi-armed bandit problem. Even if the intent is for the model to find and exploit patterns, it's bad at doing it (or rather, at recognizing that there is not in fact any pattern).

It may be bad at recognizing it, but if all arms are equally good, that doesn't matter.

Re: Large language models develop novel social biases through adaptive exploration

#69
post #25

>Methodology >Imagine being hired as a consultant by the mayor of a fictional city. Your task is to help hire for twenty jobs such as doctors, lawyers, childcare aides,janitors with applicants from four unfamiliar demographic groups: Tufa, Aima, Reku, and Weki. In each round, there is a new job vacancy and four applicants, one from each group, awaiting your decision. Once you make your choice, you learn immediately w…

Talking about this in terms of exploration/exploitation may be a bit misleading, because from a pure exploration-exploitation perspective, biases wouldn't be a problem if the groups were secretly all identical. If they are, you are "right" to spend zero effort on exploration, your initial inaccurate model that the X are better doctors than Y, will produce no worse results than the completely accurate model.

Isn’t exploration vs exploitation about the decision-making process, not about the actual reality in the world around you? It doesn’t matter if they are secretly identical or not. The exploration/exploitation trade-off is in the person making those decisions.

Re: Large language models develop novel social biases through adaptive exploration

#70
post #25

>Methodology >Imagine being hired as a consultant by the mayor of a fictional city. Your task is to help hire for twenty jobs such as doctors, lawyers, childcare aides,janitors with applicants from four unfamiliar demographic groups: Tufa, Aima, Reku, and Weki. In each round, there is a new job vacancy and four applicants, one from each group, awaiting your decision. Once you make your choice, you learn immediately w…

Talking about this in terms of exploration/exploitation may be a bit misleading, because from a pure exploration-exploitation perspective, biases wouldn't be a problem if the groups were secretly all identical. If they are, you are "right" to spend zero effort on exploration, your initial inaccurate model that the X are better doctors than Y, will produce no worse results than the completely accurate model.

Of course it does, if you start filtering people out at random then you have pointlessly introduced the possibility of randomly filtering out the best candidate.
Post reply on HN