Live data from Hacker News

Large language models develop novel social biases through adaptive exploration

openreview.net

121–124 of 124 posts

Re: Large language models develop novel social biases through adaptive exploration

#121
post #27

Earlier quoted context omitted.

But these scenarios are obviously ambiguous nonsense, which an LLM will pick up on. And given to the lack of training data on such scenarios, surely the activations are mostly random noise? It seems much more interesting to look for biases that appear robustly across different realistic scenarios that would actually be influenced by the training data

My comment is literally explaining the result of the paper, in which it is shown that LLMs can and do develop biases based on text appearing in their training data set even where such text is not in any training example connected with a systematically more positive or systematically more negative outcome. In other words, if the text "X is wet" and the text "Y is wet" and the text "X is dry" and the text "Y is dry" ea…

I'm also making a statistical observation. Saying a model "picks up on" a concept is standard shorthand, same as saying it has "learned.” What I meant is that the model has trained on plenty of neutral proper nouns that have negligible influence on the distribution of the following tokens, so the model is already conditioned towards treating them neutrally.

Not perfectly neutrally, as you said. But by your definition the only "unbiased" model is one whose output distribution perfectly matches the training distribution, i.e. one that memorized it. All LLMs have some amount of "bias” on literally every possible input.

The tribe names are no different. In the paper they run the same game again, and the bias is different every time. There's no innate preference between them trained into the model, just noise that's revealed due to a lack of any other signal. In a real situation with actually relevant information about the candidate in context, that noise is drowned out.

The more interesting thing to look for would be a bias that's strong enough to persist across different contexts. For example, is "banananow" consistently followed by positive tokens more than "pearian" across a diverse set of realistic prompts, by enough that someone could actually exploit it? The paper shows that’s explicitly not the case for made up tribe names.

What it does show, from what I can gather, is that bias can form inside a feedback loop. The model gets a success or failure result after each hire, and if a hire from one tribe happens to fail early on, the model steers that tribe away from that job for the rest of the game, even though every candidate had the same odds.

Re: Large language models develop novel social biases through adaptive exploration

#122
post #111

Earlier quoted context omitted.

> This seems to be one of the core alignment problems to me. See also: Gastown, the agent management project that could only end up working on Gastown, unceremoniously and quietly set aside. I did not know this! Any link to an announcement or autopsy of sorts (even if not by the initiator of that project)? I mean, it was pretty expensive, wasn't it? A few tens of thousands of dollars, IIRC?

This is the closest I could find to a post mortem from the creator: https://yegge.ai/essays/the-shape-of-things-to-come/ But the GasTown part is barely a single paragraph that I could not make sense of. Like, what’s the Opus “tic”? Why was it so fatal to GasTown? As someone who only ever accessed Anthropic models through other harnesses like Copilot, I have no idea. I do think what he’s saying roughly resembles what…

> Like, what’s the Opus “tic”?

The article's description gives some idea, but I presume you saw that: "the 'just two more things' tic, which prevented Opus from ever converging on being ready to do real work—it always wanted to fiddle with Gas Town itself."

Sounds like he ran into automated yak shaving. But it doesn't really explain why he couldn't fix it at the harness level.

Speculating, if you're trying to build something automatic, then you want constrained responses from each task you assign the model, otherwise you can get an endless explosion of work. The "Change 'Add to Cart' to blue" challenge parodies this: https://opusfived.dev/

Tangentially, reading the rest of that post gives me the impression that the author might benefit from an intervention. "AI psychosis" seems like it could be a relevant label here.

Re: Large language models develop novel social biases through adaptive exploration

#123
post #76

Earlier quoted context omitted.

I don't understand what you suggest that implies?

I think they're saying that while it doesn't matter, the agent and human "do not actutally know" that it does not matter. Philosophy sometimes says that knowledge is a "justified true belief"*; in this experiment, agents and humans have incorrectly justified a false belief that some applicants are better for certain roles. * other times, it says this isn't good enough

Re your asterisk, the justified true belief (JTB) criteria are considered necessary, but no longer considered sufficient as a definition of knowledge.

Because of that, JTB is often treated as a useful first approximation.

It's easy to see why the JTB criteria are necessary:

Belief: if you don't believe it then you can't count it as knowledge.

Truth: a belief is not (valid) knowledge if it's false.

Justification: accidentally getting the right answer isn't normally knowledge.

But the original claim for JTB was that it was a sufficient definition of knowledge. Later critiques like Gettier's showed that this is not generally true, i.e. there are edge cases for which it fails. In many scenarios, those edge cases don't matter much. So you end up with JTB being an imperfect but useful definition.

Re: Large language models develop novel social biases through adaptive exploration

#124

Earlier quoted context omitted.

I think it's less obtuse to make the point that there are social-prejudicial structures encoded in the latent space. "associative pattern seeking logics" is a bit incoherent, and distributional semantics doesn't live at a level accessible to cultural analysis and theoretics, IE film critique. If you walked into a film theory class and posited that you could derive every single encoded interpretation of a film by memo…

> the anti-distributional semantics crowd Who is this crowd specifically? The stochastic parrots people? Noam Chomsky? I don't think they're good representatives of media theory at all whatsoever. The humanities are much more diverse than they're made out to be in this crap AI culture war. > distributional semantics doesn't live at a level accessible to cultural analysis and theoretics They might not have computation…

> Who is this crowd specifically? The stochastic parrots people? Noam Chomsky?

I'm thinking more the Noam Chomsky and John Searle variety. For the stochastic parrots people, which I assume you to mean the no-skin-in-the-game bloggers, I'm not concerned about them. I find a lot of the rhetoric around LLMs to be eye-roll worthy, most people slinging it often lack a coherent theory of semantics to begin with, let alone an understanding of logical induction/statistics. Take the definitional entailment that these models are ampliative. This small fact undermines quite a lot of the naive mental model people have of what on transformers even are. No intentional theory, no position driving the argument, the opposition isn't substantial. The virtue signal is valid, but from strangers is uninteresting.

As you mention the humanities is very diverse, linguistics is no exception. It's not that nobody was on the corner of distributional semantics, but it's been a long road and for much of its life results were routinely dismissed in the mainstream out of dogmatism. Noam Chomsky is a good pull because I believe he's the most prominent example of this chauvinism. Vague hand-waives about explanatory power, etc.

> I'm thinking of post-structuralists especially.

I understand what you mean. It's probably not coincidence that Chomsky is a vocal critic of the school. I think many different post-modernist camps even beyond post-structuralism actually are amenable to the implications of distributional semantics, and are probably the better equipped for it. I think there are nuanced problems unifying the two, but it's a work day so I won't elaborate.

> It would depend a lot on the specific school and instructor but in general I really don't think you would. I've taken classes like this and people were far more open minded and critical than you might assume.

I was just being cute with that, really. Although in this case, I think critique is itself the problematic lever, acceptance is more an exercise of apologetics, which is the weak point of post-structuralism in many ways.

> By "associative logics" I hand-wavingly mean all manner of categorical thinking, "amygdala" thinking, mapping, putting things into buckets, hedging.

I assumed so, but it's more or less a long phrase to repeat the same concept. Relations (IE logic), mappings, morphisms, associations, patterns, etc. Same-side, same-coin. Informal, formal, take a position of drawing no line here and the shadows scatter. The extraneous qualifiers aren't what confused me though, it's that gerund "seeking". It sticks out enough to imply additional structure, but one that isn't contextually indicated. Taken in a conservative form, I considered pattern-seeking to be interchangeable with pattern-constructing, which then returns it to just being an extraneous qualifier (hence the "bit incoherent", it's a mild interpretive non-confidence).

Post reply on HN