Live data from Hacker News

Emotion concepts and their function in a large language model

anthropic.com

51–60 of 212 posts

Re: Emotion concepts and their function in a large language model

#51

> Note that none of this tells us whether language models actually feel anything or have subjective experiences. You’ll never find that in the human brain either. There’s the machinery of neural correlates to experience, we never see the experience itself. That’s likely because the distinction is vacuous: they’re the same thing.

LLMs are disembodied and exist outside of time.

Bundle of tokens comes in, bundle of tokens comes out. If there is any trace of consciousness or subjectivity in there, it exists only while matrices are being multiplied.

Re: Emotion concepts and their function in a large language model

#52

> Note that none of this tells us whether language models actually feel anything or have subjective experiences. You’ll never find that in the human brain either. There’s the machinery of neural correlates to experience, we never see the experience itself. That’s likely because the distinction is vacuous: they’re the same thing.

LLMs are disembodied and exist outside of time. Bundle of tokens comes in, bundle of tokens comes out. If there is any trace of consciousness or subjectivity in there, it exists only while matrices are being multiplied.

That’s true by definition. They’re only on when they’re on. Are you making a broader point that I’m missing?

Re: Emotion concepts and their function in a large language model

#53

There was a really old project from mit called conceptnet that I worked with many years ago. It was basically a graph of concepts (not exactly but close enough) and emotions came into it too just as part of the concepts. For example a cake concept is close to a birthday concept is close to a happy feeling. What was funny though is that it was trained by MIT students so you had the concept of getting a good grade on a…

> the concept of getting a good grade on a test as a happier concept than kissing a girl for the first time.

Were the concepts weighted by response counts? I’d imagine a good grade is a happy concept for everyone, but kissing a girl for the first time might only be good for about 50% of people.

Re: Emotion concepts and their function in a large language model

#54
> Since these representations appear to be largely inherited from training data, the composition of that data has downstream effects on the model’s emotional architecture. Curating pretraining datasets to include models of healthy patterns of emotional regulation—resilience under pressure, composed empathy, warmth while maintaining appropriate boundaries—could influence these representations, and their impact on behavior, at their source.

What better source of healthy patterns of emotional regulation than, uhhh, Reddit?

Re: Emotion concepts and their function in a large language model

#55

There was a really old project from mit called conceptnet that I worked with many years ago. It was basically a graph of concepts (not exactly but close enough) and emotions came into it too just as part of the concepts. For example a cake concept is close to a birthday concept is close to a happy feeling. What was funny though is that it was trained by MIT students so you had the concept of getting a good grade on a…

Were there published results from the project?

Re: Emotion concepts and their function in a large language model

#56
post #23

>... emotion-related representations that shape its behavior. These specific patterns of artificial “neurons” which activate in situations—and promote behaviors—that the model has learned to associate with the concept of a particular emotion. .... In contexts where you might expect a certain emotion to arise for a human, the corresponding representations are active. >For instance, to ensure that AI models are safe an…

> Force-set to 0, "mask"/deactivate those representations associated with bad/dangerous emotions. Neural Prozac/lobotomy so to speak.

More complex than that, but more capable than you might imagine: I’ve been looking into emotion space in LLMs a little and it appears we might be able to cleanly do “emotional surgery” on LLM by way of steering with emotional geometries

Re: Emotion concepts and their function in a large language model

#57

So should I go pursue a degree in psychology and become a datacenter on-call therapist?

Hah, I have been thinking about trying to study LLM psychology, nice to see that Anthropic is taking it seriously, because the mathematical psychology tools that can be invented here are going to be stunning, I suspect.

Imagine coding up a brand new type of filter that is driven by computational psychology and validated interventions, etc

Re: Emotion concepts and their function in a large language model

#58

So should I go pursue a degree in psychology and become a datacenter on-call therapist?

It's still too early to tell, but it might make sense at some point. If because of symmetry and universality we decide that llms are a protected class, but we also need to configure individual neurons, that configuration must be done by a specialist.

It might simply reduce down to a big batch of sliders and filters no different than a fancy audio equalizer: Anthropic was operating on neurons in bulk using steering vectors, essentially, as I understand it.

Re: Emotion concepts and their function in a large language model

#59

Something they don’t seem to mention in the article: Does greater model “enjoyment” of a task correspond to higher benchmark performance? E.g. if you steer it to enjoy solving difficult programming tasks, does it produce better solutions?

Pretty easy to test, I’d imagine, on a local LLM that exposes internals.

I’d suspect that the signals for enjoyment being injected in would lead towards not necessarily better but “different” solutions.

Right now I’m thinking of it in terms of increasing the chances that the LLM will decide to invest further effort in any given task.

Performance enhancement through emotional steering definitely seems in the cards, but it might show up mostly through reducing emotionally-induced error categories rather than generic “higher benchmark performance”.

If someone came along and pissed you off while you were working, you’d react differently than if someone came along and encouraged you while you were working, right?

Re: Emotion concepts and their function in a large language model

#60
post #40

> Note that none of this tells us whether language models actually feel anything or have subjective experiences. You’ll never find that in the human brain either. There’s the machinery of neural correlates to experience, we never see the experience itself. That’s likely because the distinction is vacuous: they’re the same thing.

See also: Functionalism [1]. [1] https://en.wikipedia.org/wiki/Functionalism_%28philosophy_of...

See also: Process Philosphy [0]

[0] https://plato.stanford.edu/entries/process-philosophy/

Post reply on HN