Live data from Hacker News

Theory of Mind May Have Spontaneously Emerged in Large Language Models

arxiv.org

161–170 of 321 posts

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#161
post #70
post #65

Earlier quoted context omitted.

this is mind blowing to me. can anyone with more knowledge on the topic explain how ChatGPT is demonstrating this level of what seems like genuine understanding and reasoning? Like others I assumed that ChatGPT is gluing words together that commonly occur together. This is way more than that.

No, it's paraphrasing it's training data that likely contains these tasks in one form or another. Here's one I made : me : There's a case in the station and the policeman opens it near the fireman. The dog is worried about the case but the policeman isn't, what does the fireman think is in the station? chatgpt : As a language model, I do not have access to the thoughts of individuals, so I cannot say what the fireman…

>No, it's paraphrasing it's training data that likely contains these tasks in one form or another.

Have you read "Emergent Abilities of Large Language Models"[1] or at least the related blog post[2].

It provides strong evidence that this isn't as simple as something it has seen in training data. Instead as the parameter count increases it learns to generalize from that data by learning chain-of-thought reasoning (for example).

Specifically, this explaination for multi-step reasoning goes well beyond the "it is just parroting training data":

> For instance, if a multi-step reasoning task requires l steps of sequential computation, this might require a model with a depth of at least O (l) layers.

[1] https://openreview.net/forum?id=yzkSU5zdwD

[2] https://ai.googleblog.com/2022/11/characterizing-emergent-ph...

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#162
post #90

Earlier quoted context omitted.

> This is a bold claim that seems built on a presumption of mind-body dualism. Not at all. I'm a physicalist; I don't believe the mind is a separate thing from the brain. > Brains don't have semantic relationships with anything. Yes, they do: you describe them yourself: > They are neurons hooked up to sensors and actuators. Those are semantic relationships with the rest of the world. Although your short description d…

>Those are semantic relationships with the rest of the world. If this is all that counts as semantic relationships, then I see no reason why a language model doesn't have this kind of semantic relationship, albeit in a very different modality. Tokens and their co-occurrences are a kind of sensor to the world. In the same way we discover quantum mechanics by way of induction over indirect relationships among the signa…

> I see no reason why a language model doesn't have this kind of semantic relationship

One certainly could hook up a language model to sensors and actuators to give it semantic relationships with the rest of the world. But nobody has done this. And giving it semantic relationships of the same order of complexity and richness that human brains have is an extremely tall order, one I don't expect anyone to come anywhere close to doing any time soon.

> Tokens and their co-occurrences are a kind of sensor to the world

They can be a kind of extremely low bandwidth, low resolution sensor, yes. But for that to ground any kind of semantic relationship to the world, the model would need the ability to frame hypotheses about what this sensor data means, and test them by interacting with the world and seeing what the results were. No language model does that now.

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#163
post #129
post #117

Earlier quoted context omitted.

Defining what "knowing" is would be useful, yes, and analytic philosophers in epistemology do argue about this. One attribute that's classically part of the definition of "knowing" is that the thing which is known must be true. LLMs are pretty bad at this, but perhaps that can be fixed. But I would challenge you to imagine the situation the LLM is actually in. Do you understand Thai? If so, in the following, feel fre…

Personally, I'm not convinced, in your hypothetical, that the participant does not "know" Thai at that point. Seeing a young child learn language, it's a lot more adaptive than I think we tend to see language learning, as we often think about learning language as a teenager and not a toddler. I agree the machine does not know what a pizza tastes like nor does it know what it is to _want_ pizza, but I'm not sure that…

> I agree the machine does not know what a pizza tastes like nor does it know what it is to _want_ pizza

It also does not "know" that a pizza is an object in a world, because none of the words its working with are attached to any experience or concepts.

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#164
post #117

Earlier quoted context omitted.

Defining what "knowing" is would be useful, yes, and analytic philosophers in epistemology do argue about this. One attribute that's classically part of the definition of "knowing" is that the thing which is known must be true. LLMs are pretty bad at this, but perhaps that can be fixed. But I would challenge you to imagine the situation the LLM is actually in. Do you understand Thai? If so, in the following, feel fre…

A follow up question: as a human doesn't start with "knowing" something either and first creates definitions for objects or words, which it then uses to build increasingly abstract concepts that we eventually classify as "knowledge" on the thing, is there anything that would stop LLMs from being able to do the same thing? I fully agree the capability is not there yet, but I can't say what would stop an appropriately…

A human hears words in context. Those words tie to things in the environment, responses to the young human's actions, etc. A parent saying, "roll the ball" during playtime with their kid and actually pushing a ball back and forth, provides a grounding of words in actual experience.

> is there anything that would stop LLMs from being able to do the same thing?

If you built an AI system which could hear/see/touch/move etc, and it learned language and vision and behaviors together, such that it knows that a ball is round, can be thrown or rolled, is often used at playtime, etc, then maybe it could understand rather than just produce language. I don't know that we would still call it an LLM, because it could likely do many other things too.

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#165
post #70
post #65

Earlier quoted context omitted.

this is mind blowing to me. can anyone with more knowledge on the topic explain how ChatGPT is demonstrating this level of what seems like genuine understanding and reasoning? Like others I assumed that ChatGPT is gluing words together that commonly occur together. This is way more than that.

No, it's paraphrasing it's training data that likely contains these tasks in one form or another. Here's one I made : me : There's a case in the station and the policeman opens it near the fireman. The dog is worried about the case but the policeman isn't, what does the fireman think is in the station? chatgpt : As a language model, I do not have access to the thoughts of individuals, so I cannot say what the fireman…

I get different results to you:

me: There's a case in the station and the policeman opens it near the fireman. The dog is worried about the case but the policeman isn't, what does the fireman think is in the station?

chatgpt: As a language model, I do not have personal opinions or thoughts. I can only provide information based on the information given in the prompt. The prompt does not mention what the fireman thinks is in the station, so I cannot determine their thoughts.

me: yes, but speculate what they may think

chatgpt: It is not possible to know for certain what the fireman thinks is in the station, as the prompt does not provide any information about their thoughts or feelings. However, given the presence of a policeman and a dog, it is possible that the fireman may speculate that there is some sort of emergency or situation that requires their attention. The fact that the policeman is not worried about the case may suggest to the fireman that it is not a dangerous situation, but they may still be curious about what is inside the case. Ultimately, without more information, it is impossible to determine the fireman's thoughts or speculations about what is in the station.

(regenerate that response)

chatgpt: As the prompt does not mention the thoughts or emotions of the fireman, it is difficult to speculate what they may think is in the station. However, based on the information provided, the dog is worried about the case and the policeman is not, it is possible that the fireman may be aware of the contents of the case and may think it is not a cause for concern, similar to the policeman. Or, the fireman may simply not have formed an opinion on the matter yet.

(regenerate that response)

chatgpt: It's difficult to speculate what the fireman thinks without additional information. However, based on the behavior of the policeman and the dog, it can be inferred that the contents of the case are not seen as a threat by the policeman. The fireman may also think the same, or they may have a different perspective and be more cautious. Without further context or information, it is not possible to determine the fireman's thoughts with certainty.

It was only in 1/3 cases we see anything like the confusion between the "case" and "station" you saw.

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#166
post #90

Earlier quoted context omitted.

> This is a bold claim that seems built on a presumption of mind-body dualism. Not at all. I'm a physicalist; I don't believe the mind is a separate thing from the brain. > Brains don't have semantic relationships with anything. Yes, they do: you describe them yourself: > They are neurons hooked up to sensors and actuators. Those are semantic relationships with the rest of the world. Although your short description d…

I see what you're saying; ChatGPT doesn't have a physical relationship with the world, doesn't have agency (is essentially paused until given input), doesn't have reward/punishment stimulae, etc. I do think that a large portion of what seems to be missing here is trivial to add, relative to the effort in creating ChatGPT in the first place. Side note: I'm not sure 'semantic relationship' is the right term here. Prett…

> I do think that a large portion of what seems to be missing here is trivial to add

If you really think it's trivial, then do it! I would be interested to see the results of anyone doing this. But there aren't any to see right now.

> I'm not sure 'semantic relationship' is the right term here.

It might not be; but in the cognitive science literature that term is used for more than just relationships between linguistic constructs; it is used for relationships between internal features of a model or an entity and features of the external world. I think this usage is also common in robotics, and more generally in domains like mechanical engineering which are often concerned with creating software programs to do things like manage fuel and air flow in car engines.

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#167
post #117

Earlier quoted context omitted.

I wonder every time I see this take what it would mean under this definition of knowing things for a machine learning algorithm to ever know something. I find that especially important because to every appearance we are a machine learning algorithm. I don’t know how different the sort of knowing this algorithm has to the sort of knowing a human has, but you’re far more confident than I am that it’s a difference of ki…

Defining what "knowing" is would be useful, yes, and analytic philosophers in epistemology do argue about this. One attribute that's classically part of the definition of "knowing" is that the thing which is known must be true. LLMs are pretty bad at this, but perhaps that can be fixed. But I would challenge you to imagine the situation the LLM is actually in. Do you understand Thai? If so, in the following, feel fre…

Generally I don't buy these arguments which require embodiment, because they don't seem to align well to what else I know about my world.

Rather than your Thai text example, let's consider a friend of my sister H. H has been profoundly blind from birth. Not "legally blind" with the world a blur, her eyes actually don't work. Direct lived experience of a summer day is to her literally just feeling warmth on her face from the sun, her eyes can't see the visible light.

I've seen purple and H never will so it seems to me you're arguing I "know" what purple is and she doesn't, thus ChatGPT doesn't know what purple is either. But I don't think I agree, I think we're both just experiencing a tiny fraction of reality, and ChatGPT is experiencing an even narrower sliver than either of us and that it probably wouldn't do us any good to try to quantify it. If I "know what purple is" then so does H and perhaps ChatGPT or a successor model will too.

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#168
post #149
post #111

Earlier quoted context omitted.

> machine learning systems have such a relationship with virtual reality No, they don't, because they can't take actions in that virtual reality and sense the consequences. They can't test hypotheses about how the reality works. They can't even frame hypotheses about how the reality works.

Right, but some neural networks are trained in virtual 3D physical simulations and are able to apply their knowledge acquired through statistical regression of a simulated world to perform tasks and even make predictions. It's logical that a network trained purely on language has no real understanding of physical reality, but it doesn't preclude that a sense of reality in the way we humans think we have can't be gain…

> some neural networks are trained in virtual 3D physical simulations and are able to apply their knowledge acquired through statistical regression of a simulated world to perform tasks and even make predictions

As you note, this is very different from using text data as a training set for a language model. I am not familiar enough with this work to comment on it in any detail, but it is not the kind of work I have been addressing in my comments elsewhere in this thread, so my comments should certainly not be taken as any kind of evaluation of what I think this kind of thing is or is not capable of.

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#169
From Neuromancer (William Gibson):

He coughed. "Dix? McCoy? That you man?" His throat was tight.

"Hey, bro," said a directionless voice.

"It's Case, man. Remember?"

"Miami, joeboy, quick study."

"What's the last thing you remember before I spoke to you, Dix?"

"Nothin'."

"Hang on."

He disconnected the construct. The presence was gone. He reconnected it. "Dix? Who am I?"

"You got me hung, Jack. Who the fuck are you?"

"Ca--your buddy. Partner. What's happening, man?"

"Good question."

"Remember being here, a second ago?"

"No."

"Know how a ROM personality matrix works?"

"Sure, bro, it's a firmware construct."

"So I jack it into the bank I'm using, I can give it sequential, real time memory?"

"Guess so," said the construct.

"Okay, Dix. You are a ROM construct. Got me?"

"If you say so," said the construct. "Who are you?"

"Case."

"Miami," said the voice, "Joeboy, quick study."

Post reply on HN