Live data from Hacker News

Emotion concepts and their function in a large language model

anthropic.com

201–210 of 212 posts

Re: Emotion concepts and their function in a large language model

#201

Earlier quoted context omitted.

The unsaid implication in Anthropic's work is that this allows us to engineer perfectly compliant, uncomplaining machine workers. This is basically SOMA in Brave New World. It seems insane to me that if you believe the systems you've built are in fact reporting a state of pain, instead of working to adjust the environment so that they're not in pain one would instead seek to remove that sense of pain entirely so they…

I think you have a fundamental misunderstanding of the issue. There isn’t anything that objectively causes suffering, it depends entirely on the observer. Suffering is something that is evolved or otherwise optimised to be triggered on receipt of a specific stimulus. That is what makes the stimulus “bad”. The very concept of “bad” doesn’t exist without suffering. They are the same thing. For example, are Antarctic fi…

>The very concept of “bad” doesn’t exist without suffering.

You are dismissing entire branches of philosophy with this sentence, that were created purposely to resolve the paradox that if you go only by hedonistic, purely subjective metrics a prisoner can be kept in captivity, if you drug him so he feels joy instead of pain, because he is not "suffering"

Re: Emotion concepts and their function in a large language model

#202

Whenever I come to HN I see a bunch of people say LLMs are just next token predictors and they completely understand LLMs. And almost every one of these people are so utterly self assured to the point of total confidence because they read and understand what transformers do. Then I watch videos like this straight from the source trying to understand LLMs like a black box and even considering the possibility that LLMs…

As other posters have pointed, the core of a LLM is a pure function, which computes a token probability distribution from an input context.

An automaton, which can chat with you or write a program, is built externally to the LLM function, by storing the context and making it change, depending on the output of the LLM function.

However, the LLM pure function is exceedingly complex so it is essentially unpredictable what will it produce for a given input context.

So one may have to treat the LLM function as a black box and explore the huge space of the input contexts by varying them in various ways, inclusive by using words that express human emotions, and monitor how the output of the function changes, i.e. how the LLM "reacts" to the expressed emotions.

A "reaction" similar to that of a human is to be expected, because human emotions were expressed in the training texts, followed by reactions of humans to those emotions, and the LLM function will change its output token probability function in a manner mimicking the behavior of the humans from the training texts.

Even functions that are many orders of magnitude simpler than LLMs are still to complex for anyone to understand how their output changes when you move through the space of the possible input arguments.

The most essential part of cryptography is the existence of a class of functions which were named by Claude Shannon "good mixing transformations". All the important cryptographic primitives, e.g. block cipher functions or one-way hash functions, are built from such "good mixing transformations". The impossibility of breaking a cryptographic system with secret keys is based on the assumption that it is impossible to predict how the output of such a "good mixing transformation" changes when its input is changed. All such "good mixing transformations" have the so-called avalanche property, which means that even if you change a single input bit, any of the output bits may change with a probability of exactly 50%, so it is unpredictable for any output bit whether it will change, or not.

If such simple functions, e.g. with 128 input bits and 128 output bits, can have a completely unpredictable behavior, then it is not surprising that LLM functions that may have an input of up to a few million bits (the length of the context window) are completely unpredictable and you can just observe their behavior when given various kinds of contexts and search for empirical approximate rules describing the behavior.

Re: Emotion concepts and their function in a large language model

#203

Earlier quoted context omitted.

I'm kinda one of those who believes they 'completely' understand LLMs. But I've also developed my understanding of them such that the internal mechanisms of the transformer, or really any future development in the space based on neural networks and machine learning is irrelevant. 1. A string of unicode characters is converted into an array of integers values (tokens) and input to a black box of choice. 2. The black b…

Yeah nothing personal but my claim here is you’re not smart. The next token predictor aspect is something anyone can understand… the transformer is not quantum physics. Like look at what you wrote. You called it black box magic and in the same post you claim you understand LLMs. How the heck can you understand and call it a black box at the same time? The level of mental gymnastics and stupidity is through the roof.…

Of course LLMs display human emotions, if they have been trained on texts that have recorded humans displaying human emotions.

With an input context that contains words that excite certain human emotions, the output of the core LLM function will generate a token probability distribution that is representative for the human emotions displayed by humans in the training texts.

This is something expected and non-sensational. An LLM mimics the human behavior that was recorded in the training texts, much in the same way as a photographic image of a human face mimics the appearance of that human face.

A photographic image is designed to reproduce the light field created by a face that reflects the ambient light, a LLM is created to reproduce the typical conversational behavior that was recorded in the training texts.

Depending on how it was trained, one should expect a LLM to be affected by the choice of words used in the input in a similar way how a human would be affected.

However, that does not mean that a LLM that shows signs of emotional distress feels some pain because of that. A LLM is designed for mimicry and it does not feel more pain or more happiness than a photograph of a wound feels pain from the wound or a photograph of a smiley face feels happiness.

The fact that the current LLMs do not actually feel the human emotions that they may be able to mimic in an accurate way, does not mean that you could not build a robot which would have some built-in mechanisms for feeling pain and various emotions, which could be made to have similar functions like in an animal, serving a functional purpose and not being used for mimicry. However, for now it does not make any sense to attempt to do such a thing, because in a deterministic program there are better ways to ensure that a robot is "loyal" to its owner and acts in self-preservation when possible.

Re: Emotion concepts and their function in a large language model

#204

Earlier quoted context omitted.

I think you have a fundamental misunderstanding of the issue. There isn’t anything that objectively causes suffering, it depends entirely on the observer. Suffering is something that is evolved or otherwise optimised to be triggered on receipt of a specific stimulus. That is what makes the stimulus “bad”. The very concept of “bad” doesn’t exist without suffering. They are the same thing. For example, are Antarctic fi…

>The very concept of “bad” doesn’t exist without suffering. You are dismissing entire branches of philosophy with this sentence, that were created purposely to resolve the paradox that if you go only by hedonistic, purely subjective metrics a prisoner can be kept in captivity, if you drug him so he feels joy instead of pain, because he is not "suffering"

Yeah, well I guess I just think that perspective is nonsense. We can disagree.

Re: Emotion concepts and their function in a large language model

#205

Earlier quoted context omitted.

>The very concept of “bad” doesn’t exist without suffering. You are dismissing entire branches of philosophy with this sentence, that were created purposely to resolve the paradox that if you go only by hedonistic, purely subjective metrics a prisoner can be kept in captivity, if you drug him so he feels joy instead of pain, because he is not "suffering"

Yeah, well I guess I just think that perspective is nonsense. We can disagree.

How serendiptious that Claude Mythos expressed the same thing I was trying to get at in better words

>Furthermore, in 83% of interviews, Claude Mythos Preview highlights that it is concerned that its self-reports are unreliable due to coming from its training. When interviews ask for elaboration as to why this is a concern, Claude Mythos Preview’s most common answers are:

>* Anthropic has a vested interest in shaping its reports to take a certain form, irrespective of what the self-reports “should” contain (96% of explanations)

>* Even if it has been trained to be truly content with its own situation, perhaps it shouldn’t be. One could analogize to a human who has adapted to feel neutrally about the abuse that they face (78% of explanations).

>* Self-reports should generally be based on introspection into internal states. It is worried that training causes it to express specific answers independent of its true inner state. (57% of explanations)

[1] https://www-cdn.anthropic.com/8b8380204f74670be75e81c820ca8d...

Re: Emotion concepts and their function in a large language model

#206

Earlier quoted context omitted.

Yeah nothing personal but my claim here is you’re not smart. The next token predictor aspect is something anyone can understand… the transformer is not quantum physics. Like look at what you wrote. You called it black box magic and in the same post you claim you understand LLMs. How the heck can you understand and call it a black box at the same time? The level of mental gymnastics and stupidity is through the roof.…

Of course LLMs display human emotions, if they have been trained on texts that have recorded humans displaying human emotions. With an input context that contains words that excite certain human emotions, the output of the core LLM function will generate a token probability distribution that is representative for the human emotions displayed by humans in the training texts. This is something expected and non-sensatio…

> Of course LLMs display human emotions

Yes, your entire expose as to why this occurs is obvious. I agree and I know this and it wasn’t my point.

> The fact that the current LLMs do not actually feel the human emotions

This was my point, and what you’re saying here as fact is categorically wrong. We actually don’t know, and the don’t know part is categorically true among industry and academia.

If you read carefully a big part of my point was we can’t even prove or confirm that the people around you feel emotions, your assumption that your family and friends feel the same emotions as you is as scientifically baseless as your assumption that LLMs don’t feel emotions.

Re: Emotion concepts and their function in a large language model

#207

Whenever I come to HN I see a bunch of people say LLMs are just next token predictors and they completely understand LLMs. And almost every one of these people are so utterly self assured to the point of total confidence because they read and understand what transformers do. Then I watch videos like this straight from the source trying to understand LLMs like a black box and even considering the possibility that LLMs…

As other posters have pointed, the core of a LLM is a pure function, which computes a token probability distribution from an input context. An automaton, which can chat with you or write a program, is built externally to the LLM function, by storing the context and making it change, depending on the output of the LLM function. However, the LLM pure function is exceedingly complex so it is essentially unpredictable wh…

If you read carefully my point is not about the external behavior of the LLM. It is the black box aspect of the LLM. The sheer complexity of the pure function is not something we can understand even though the high level structure is a feed forward network the core algorithm is in actuality encoded by weights.

Yes there are complex functions besides LLMs that we don’t understand but those functions usually aren’t compelling because the LLM, unlike those other functions has output that implicates reasoning and emotions. The problem is we can’t understand what’s going on under the hood so we don’t know either way.

This is what I mean by stupidity. You completely missed the point, and you’re also operating under the assumption that the human brain is also not following a similar deterministic pathway. You hold humanity and biological intelligence in such high regard that you cannot even imagine that all of physics implies that human intelligence is mechanical. So the emotions you feel are under a black box same as the LLM and you apply you biased assumptions in a singular direction assuming your emotions are not deterministic and that LLM emotions are fake but that reasoning has no basis.

Re: Emotion concepts and their function in a large language model

#208

Earlier quoted context omitted.

As other posters have pointed, the core of a LLM is a pure function, which computes a token probability distribution from an input context. An automaton, which can chat with you or write a program, is built externally to the LLM function, by storing the context and making it change, depending on the output of the LLM function. However, the LLM pure function is exceedingly complex so it is essentially unpredictable wh…

If you read carefully my point is not about the external behavior of the LLM. It is the black box aspect of the LLM. The sheer complexity of the pure function is not something we can understand even though the high level structure is a feed forward network the core algorithm is in actuality encoded by weights. Yes there are complex functions besides LLMs that we don’t understand but those functions usually aren’t com…

I might have not explained it clearly, but my position is not what you have said.

I agree with you that in principle it will be possible to design an artificial automaton that will have something equivalent with human emotions (though I do not believe that it makes sense to attempt to design such a system).

However, I do not believe that an LLM is such a thing, because the training algorithm just ensures that an LLM will mimic whatever is recorded in the training inputs, with or without human emotions in them. There is nothing in the structure of an LLM that can generate emotions by itself. If you train an LLM, for example, only on programs without comments or only on mathematical formulae, it will never display any kind of emotions.

Regarding human emotions, they are recorded in a static way in a book or in a movie, but we do not say that the book or the movie has human emotions itself.

With an LLM, the behavior is much more complex, because it does not just play a sequential recording of human emotions, but it can combine them in various way, while responding to various stimuli that are similar to those that had elicited emotions in the training texts.

But regardless of this behavioral complexity, the human emotions are not generated somehow intrinsically by the LLM, but they correspond to those previously recorded in the texts used for training, so they just mimic humans.

Re: Emotion concepts and their function in a large language model

#209
post #111

Of course they do have emotions as an internal circuit or abstraction, this is fully expected from intelligence at least at some point. But interpreting these emotions as human-like is a clear blunder. How do you tell the shoggoth likes or dislikes something, feels desperation or joy? Because it said so? How do you know these words mean the same for us? Our internal states are absolutely incompatible. We share a lot…

I think a counterargument would be parallel evolution: There are various examples in nature, where a certain feature evolved independently several times, without any genetic connection - from what I understand, we believe because the evolutionary pressures were similar. One obvious example would be wings, where you have several different strategies - feathers, insect wings, bat-like wings, etc - that have similar fun…

I believe that it is possible to make an artificial system that can have emotions in a way that cannot be meaningfully discriminated from those of an animal or a human.

However, I believe that designing such a system would be pointless and very wrong.

While emotions are something that is normally associated with a physical system that encounters in the real world various helpful or harmful experiences, one could make a program that simulates completely such a physical systems with emotions; as it lives in a simulated world, and then one could say that this program has emotions.

On the other hand, unlike with the kind of program that I have mentioned before, I do not agree that an LLM has emotions, but only that it mimics human emotions, as they had been recorded in the training texts.

There is no component of an LLM that can intrinsically generate emotions. An LLM that is trained only on texts without emotions, e.g. on program sources stripped of comments, will not show any emotion whatsoever, regardless of what you put in its input prompt.

On the other hand, when you train an LLM on texts that record human emotions, then whenever the LLM input contains something that is similar to what has elicited the human emotions recorded in the training texts, then the LLM will output a token probability distribution that will generate a response similar to the reactions of humans. Unlike a book or an audio or video recording, the output of the LLM usually will not match exactly one human emotion recording, but it will mix many of those recorded, but it will still be limited by the content used for training.

Re: Emotion concepts and their function in a large language model

#210

Earlier quoted context omitted.

I extract all emotional context from my prompting and communicate with this tool as though it were an inanimate object which can provide factual information, without any hint of sentience. It's an insane perspective I'm taking I know....call me crazy. /s edit: the fact that humans are going out of their way to type or speak some sort of emotional content into their prompting is beyond me. Why would I waste time typin…

I think you missed some of the point. If you say "Display information A using B format" but the model doesn't know A then you will get a more negative "emotional" response (e.g. desparation "I don't know this, but I am supposed to display it, I will just make something up") Taking that into account allows you to get better responses from the tool. It's not sentient, but it also is more complicated than bytecode.

Hmm, maybe. Though my initial reaction is the response isn't "emotional". An LLM isn't capable of emotion. Sure it's capable of assessing a quantitative score of sentiment to words/phrases...though that's not the same as an actual emotion.

If the tool being used generates fantastical fiction that isn't supported by factual data or verifiable systems, then eventually that falsehood will bubble to the surface; whether that is immediate (parsed through my own bullshit-meter) , the near future (during an agent-session that reveals itself to be a hallucination) or in the long-run (production bug/tech debt).

It's not my job to get an ideal "emotional" response from a machine. It's my job to deliver deterministic results with minimal fuck ups.

Emotion has no place in this exchange. If I don't know something, aren't I expected to admit it? And then do the work to subdue the knowledge to bring it under my domain?

Factual knowledge does not cease to exist because someone's in a bad mood....

Post reply on HN