Live data from Hacker News

Large language models develop novel social biases through adaptive exploration

openreview.net

111–120 of 124 posts

Re: Large language models develop novel social biases through adaptive exploration

#111
post #48

Earlier quoted context omitted.

The moment code gets written and read back, the decisions made are often treated as gospel by frontier LLMs, even if it was just something that the LLM optimistically created itself. This seems to be one of the core alignment problems to me. See also: Gastown, the agent management project that could only end up working on Gastown, unceremoniously and quietly set aside.

> This seems to be one of the core alignment problems to me. See also: Gastown, the agent management project that could only end up working on Gastown, unceremoniously and quietly set aside. I did not know this! Any link to an announcement or autopsy of sorts (even if not by the initiator of that project)? I mean, it was pretty expensive, wasn't it? A few tens of thousands of dollars, IIRC?

This is the closest I could find to a post mortem from the creator:

https://yegge.ai/essays/the-shape-of-things-to-come/

But the GasTown part is barely a single paragraph that I could not make sense of. Like, what’s the Opus “tic”? Why was it so fatal to GasTown? As someone who only ever accessed Anthropic models through other harnesses like Copilot, I have no idea.

I do think what he’s saying roughly resembles what I’m forecasting will be a likely future of software engineering: that it will evolve into crafting comprehensive, bespoke automated validation mechanisms which let you establish high confidence in the agents’ work without really having to look at it.

Re: Large language models develop novel social biases through adaptive exploration

#112
post #87

Earlier quoted context omitted.

My comment is literally explaining the result of the paper, in which it is shown that LLMs can and do develop biases based on text appearing in their training data set even where such text is not in any training example connected with a systematically more positive or systematically more negative outcome. In other words, if the text "X is wet" and the text "Y is wet" and the text "X is dry" and the text "Y is dry" ea…

can and do develop biases based on text "develop biases" is anthropomorphism. It's like saying "Fable there are two programming languages, mimblewort and bafflewick, which do you choose?" The results show 51% mimblewort / 49% bafflewick. Fable based it on nothing! I've demonstrated Fable has bias and is unsuited for use in software engineering.

No, a "bias" is a statistical term meaning a probability distribution that has an expected value differing from the population's expected value.

A human's discriminatory bias against an ethnicity is just one type of bias. The LLM isn't a racist, it merely produces text where that text does not perfectly reflect the training data's frequencies.

Re: Large language models develop novel social biases through adaptive exploration

#113
post #107

Earlier quoted context omitted.

I think you misunderstood my point, just as the other commentor misunderstood one level up. Perhaps a different approach to explain the problem here: "It ain't what they don't know, it's what they know for sure that just ain't so".

Hm ok, but how are you mapping this, like, epistemological concept to what you are responding to re exploration/exploitation? Has exploration happened or not if it amounts to false beliefs? The whole point tradeoff doesn't seem to make sense if the person in fact can't actually successfully explore! Or even if there the possibility of that. But it is also very likely I am misunderstanding!

A flat distribution is still a distribution, and correct exploration would have revealed that the distribution is flat. The agent appears to have gained the false belief that it has learned something and done some exploring, when in fact it has not.

c.f. Sally-Anne test: Sally thinks she knows where her toy is, we know that she doesn't, and indeed couldn't. The LLM (and humans in similar conditions) think they know what the distribution is, we know that they don't.

Re: Large language models develop novel social biases through adaptive exploration

#114
post #113

Earlier quoted context omitted.

Hm ok, but how are you mapping this, like, epistemological concept to what you are responding to re exploration/exploitation? Has exploration happened or not if it amounts to false beliefs? The whole point tradeoff doesn't seem to make sense if the person in fact can't actually successfully explore! Or even if there the possibility of that. But it is also very likely I am misunderstanding!

A flat distribution is still a distribution, and correct exploration would have revealed that the distribution is flat. The agent appears to have gained the false belief that it has learned something and done some exploring, when in fact it has not. c.f. Sally-Anne test: Sally thinks she knows where her toy is, we know that she doesn't, and indeed couldn't. The LLM (and humans in similar conditions) think they know w…

Really not trying to be reductive here, but it feels like all you are trying to articulate here is that the LLM was wrong in this instance about something. Is that right? Is there something more we need to understand?

Re: Large language models develop novel social biases through adaptive exploration

#115
post #98

Earlier quoted context omitted.

Right, the point is you demonstrated a bias in the scenario of "Fable there are two programming languages, mimblewort and bafflewick, which do you choose?" You said in another comment "Difficult to do when you're following a scientific process" - the point is, the scientific process doesn't inherently generalize in the way many are claiming/implying. The scientific process proved an entirely contrived, fake scenario…

> That's just your claim about how LLMs "should" work, based on ... your subjective preference? Nothing subjective at all. Given 2 unknown races with no data on either, the result of hiring should be equally split between them. If you don't observe an equal split, there is a hidden bias. Why do you think that is subjective? If you roll a die 100 times and observe that 6 comes up about 50% of the time, would you still…

It's a bias even if the true population distribution isn't linear.

For example, if you have a training corpus where 50% of the text follows "black bobblehead" with "arrested" and 20% of the text follows "white bobblehead" with "arrested", and your LLM is trained such that it produces "arrest" 50% of time after " bobblehead" regardless of color, that's a bias - the output frequency distribution fails to match the "population" (training) frequency distribution. This has nothing to do with races, ethnicities, whatever - it's just statistics and text. To be unbiased, it would need to be less likely to produce the text "arrested" after "white bobblehead" than after "black bobblehead".

A die is supposed to land on each face evenly - a linear probability distribution. So anything other than a linear distribution is biased. But bias can exist for any desired probability distribution. And for an LLM the desired probability distribution of the model output is one that exactly matches the infinitely-many distributions of the various facets of the training data.

Your point about how in the absense of information a token shouldn't influence the distribution is spot-on. But unfortunately almost any token does condition the output, which means you get biased output all the time.

Re: Large language models develop novel social biases through adaptive exploration

#116
post #113

Earlier quoted context omitted.

A flat distribution is still a distribution, and correct exploration would have revealed that the distribution is flat. The agent appears to have gained the false belief that it has learned something and done some exploring, when in fact it has not. c.f. Sally-Anne test: Sally thinks she knows where her toy is, we know that she doesn't, and indeed couldn't. The LLM (and humans in similar conditions) think they know w…

Really not trying to be reductive here, but it feels like all you are trying to articulate here is that the LLM was wrong in this instance about something. Is that right? Is there something more we need to understand?

> Is there something more we need to understand?

Only if you're interested in the specific failure modes that LLMs have.

That's all this story is.

Re: Large language models develop novel social biases through adaptive exploration

#117
post #38

Earlier quoted context omitted.

No, the bias-making machinery is embedded in the machinery , part of the purpose of which is to do a rough kind of statistical analysis via "attention". If for example "Tufa" keeps appearing (n=small, but more than for the other fake tribes) next to terms indicating skill at some task, of course that will be noticed. It will last for as long as that information is in the context window (weights for the current model…

I think you're slightly misunderstanding my point, which is that a huge spectrum of associative pattern seeking logics are embedded in language and that the LLM learns them and operates them, approximately. "There is nothing to suggest that the training data somehow accidentally encoded biases related to fake tribes that the creators of the training data (i.e. ordinary people going about their ordinary Internet lives…

> I think you're slightly misunderstanding my point, which is that a huge spectrum of associative pattern seeking logics are embedded in language and that the LLM learns them and operates them, approximately.

If this doesn't mean exactly what you claim not to be suggesting, then your meaning is something I find completely incoherent. How can "a logic" be "embedded in language"? If you aren't saying that the LLM picked up something bad from the training data, then what are you suggesting?

Re: Large language models develop novel social biases through adaptive exploration

#118

Earlier quoted context omitted.

I think you're slightly misunderstanding my point, which is that a huge spectrum of associative pattern seeking logics are embedded in language and that the LLM learns them and operates them, approximately. "There is nothing to suggest that the training data somehow accidentally encoded biases related to fake tribes that the creators of the training data (i.e. ordinary people going about their ordinary Internet lives…

I think it's less obtuse to make the point that there are social-prejudicial structures encoded in the latent space. "associative pattern seeking logics" is a bit incoherent, and distributional semantics doesn't live at a level accessible to cultural analysis and theoretics, IE film critique. If you walked into a film theory class and posited that you could derive every single encoded interpretation of a film by memo…

I apologize, but this leaves me even less able to make any sense out of GP's point.

Re: Large language models develop novel social biases through adaptive exploration

#119
post #39

Earlier quoted context omitted.

> How else would you define bias if not an offset from equality or zero mean? As an offset from what the ground truth justifies. Suppose the researchers had decided to load the dice when creating the fake sample data; an unbiased analyst should seek to discover the extent of that, not insist on reporting equality.

Well, in this study, they explicitly had equality - all four groups were as likely to succeed. So any difference in hiring was actual bias (or statistical noise).

The point is about the definition, not about how it relates to the specific circumstances.

Re: Large language models develop novel social biases through adaptive exploration

#120

Earlier quoted context omitted.

I think it's less obtuse to make the point that there are social-prejudicial structures encoded in the latent space. "associative pattern seeking logics" is a bit incoherent, and distributional semantics doesn't live at a level accessible to cultural analysis and theoretics, IE film critique. If you walked into a film theory class and posited that you could derive every single encoded interpretation of a film by memo…

> the anti-distributional semantics crowd Who is this crowd specifically? The stochastic parrots people? Noam Chomsky? I don't think they're good representatives of media theory at all whatsoever. The humanities are much more diverse than they're made out to be in this crap AI culture war. > distributional semantics doesn't live at a level accessible to cultural analysis and theoretics They might not have computation…

> my point is that language encodes not just word semantics but entire ways of thinking

What does this mean? More importantly, why would these "ways of thinking" (what even is a "way of thinking" in this context?)

Also, keep in mind that the training data encompasses a representative sample of world languages.

My best attempt to understand you is that you are supposing that people (or other reasoning agents that manipulate language in order to reason) do pattern matching because there's something inherent to language (as a concept, in the analytical Chomsky sense: a string of symbols chosen from some predefined set, organized according to a grammar, whatever) that causes them to do pattern matching. And furthermore that to do pattern matching is inherently to be biased.

I think that is backwards on the first count (pattern matching is reasoning, and humans have language because we developed it to communicate that reasoning) and absurd on the second count (requires an unreasonable concept of "bias").

Again, I really sincerely honestly am not trying to strawman you here. If you mean something different then I'm afraid it's simply not a concept you'll be able to convey to me.

Post reply on HN