Live data from Hacker News

Theory of Mind May Have Spontaneously Emerged in Large Language Models

arxiv.org

271–280 of 321 posts

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#271

Earlier quoted context omitted.

People ask this question like it's meaningful... is there any proof that we are? No. Then stop asking it as if it sheds light into the similarities between humans and machines... it doesn't and it's obfuscating to that extent.

Huh. And here I thought best practice in the absence of evidence was to keep an open mind rather than asserting one extreme or the other.

Well my intuition is that the machines aren't performing comparable processes, and I'm totally open to information going the other way. I'm not open to baseless assertions otherwise.

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#272
post #166

Earlier quoted context omitted.

> I do think that a large portion of what seems to be missing here is trivial to add If you really think it's trivial, then do it! I would be interested to see the results of anyone doing this. But there aren't any to see right now. > I'm not sure 'semantic relationship' is the right term here. It might not be; but in the cognitive science literature that term is used for more than just relationships between linguist…

Look again. There are papers that hook up llms to robots with vision and other sensors. The LLM is fed descriptions of the world and then emits instructions for where to go

The thought of an LLM interacting with the world through a MUD is entertaining :)

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#273
post #250
post #244

Earlier quoted context omitted.

What would be evidence that a prediction machine had developed a theory of mind?

well, I'd first need to see that it had a Theory of Feet... ;-) More seriously, that it can actually understand and wield abstract concepts. Can it accurately and repeatedly understand that "the foot attaches to the shin bone, which attaches to the thigh bone, which attaches to the hip bone...", and that these have certain degrees of freedom, but not others, and that one foot goes in front of the other, and to easily…

"statistical re-mixer" doesn't describe these systems very well. I see this complaint a lot, that supposedly DL models can only manipulate existing content without creating anything of their own. That's just false, unless your standard for originality is so high that humans can't reach it either.

These models that have hundreds of billions of "synapses", it's not very shocking to me that they can learn the abstract form of concepts. In fact, it's kind of beautiful that human concepts have this mathematical nature. It vindicates Plato, and disappoints everyone who has claimed that language and meaning is arbitrary.

But the main issue here is that for every conceivable empirical test we can perform, you'll still make the same complaint. Even after it's demonstrated better ToM abilities than you, by predicting and explaining other people's mental states better than you can, you'll say the same thing.

Maybe it's because you think that "understanding" requires not just accuracy, but having a certain kind of inner experience that a human could relate to.

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#274
post #179

Earlier quoted context omitted.

> than you would expect from an algorithm that just predicts the next word. I think there is common mistake in this concept of just predicting the next word. While it is true that just the next word is predicted, a good way to do that is to internally imagine more than the next word and then just spit out the next word. Of course with the word after that the process repeats with a new imagination. One may say that th…

It predicts the next word based on the preceding 2000 words or so, thats the thing. And to do that takes serious modelling.

Okay. So you agree, it seems.

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#275
post #97

Earlier quoted context omitted.

OK how about this: It 'knows' language, as in it has learnt about relationships between words (and thats really underselling it, in reality it has learnt very very subtle relationships between a great many words, and it can process words about 2000 at a time (token count etc)) BUT as you say it has no outside reference, its just a bundle of weights (those weights forming models of a sort) BUT we provide the outside c…

> it wont be long before someone hooks one of these up to cameras and robot arms and teaches it to make a cup of tea or whatever. And that will at least be a start at giving these things some very simple semantic relationships with the outside world. But right now they have none.

You haven't heard of GATO then?

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#276

Earlier quoted context omitted.

I was thinking in terms of simple logic and semantics. The example I picked though muddied the waters by bringing in real-world phenomena. A better test would be anything that stays strictly within the symbolic world - the true umwelt of the language model. So, anything mathematical. After seeing countless examples of addition and documents discussing addition and procedures of addition, many order of magnitude more…

A child can 'see' maths though, they can see that if you have one apple over here and one orange over there, then you have two pieces of fruit all together. If you only ever allowed a child to read about adding, without ever being able to physically experiment with putting pieces together and counting them, likely children would not be able to add either. In fact, many teachers and schools teach children to add using…

Yet ChatGPT totally - apparently - gets 1 + 1. In fact it aces the addition table way beyond what a child or even your average adult can handle. It's only when you get to numbers in the billions that it's weaknesses become apparent. One thing it starts messing up is carry-over operations, from what I can see. Btw. the treshold used to be significanly lower, yet that doesn't convince me in the least that it's made progress in its understanding of addition. It's still just as much in the fog. And it cannot introspect and tell me what it's doing so I can point out where it's going wrong.

But I think you are right in what you are saying. Basically it not 'seeing' math as a child does, is just another way to say that it doesn't undestand math. It doesn't have a intuitive understanding of numbers. It also can't really experiment. What would experimenting mean in this context? Just more training cycles. This being math, one could have it run random sums and give it the correct answer each time. That's one way to experiment, but that wouldn't solve the issue. At some point it would reach its capacity of absorbing statistical corelations to deal with numbers large enough. It would need more neurons to progress beyond that stage.

Btw. I found this relevant article: https://bdtechtalks.com/2022/06/27/large-language-models-log...

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#277
post #273
post #250

Earlier quoted context omitted.

well, I'd first need to see that it had a Theory of Feet... ;-) More seriously, that it can actually understand and wield abstract concepts. Can it accurately and repeatedly understand that "the foot attaches to the shin bone, which attaches to the thigh bone, which attaches to the hip bone...", and that these have certain degrees of freedom, but not others, and that one foot goes in front of the other, and to easily…

"statistical re-mixer" doesn't describe these systems very well. I see this complaint a lot, that supposedly DL models can only manipulate existing content without creating anything of their own. That's just false, unless your standard for originality is so high that humans can't reach it either. These models that have hundreds of billions of "synapses", it's not very shocking to me that they can learn the abstract f…

Yes, I understand that it can appear to synthesize something new, and no, I'm not looking for some inner experience.

I'm looking for it to show an ability to wield not only a set of strings (with language associations), but something actually like the platonic ideals - objects, with properties and relations.

A few errors show quickly there is no such concept being weilded.

>> I saw a fine example of this failure the other day: "Mike's mom has four kids. three are named Danielle, Liam, and Kelly. What is the fourth kid's name?" ChatGPT's reply is explanation of how there isn't enough info in question to tell. Told "The answer is in the question.", ChatGPT just doubles down on the answer. (Sorry, couldn't find the original example)

>> "My sister was half my age when I was six years old. I'm now 60 years old. How old is my sister?" ChatGPT: "Your sister is now 30 years old". [0]

>> Or this one where ChatGPT entirely fails to understand order/sequence of events. [1]

Or a plethora of math problem fails found...

Similarly, the image "AI"s fail to understand relationships between objects (or parts of one object), and cannot abstract a particular person's image from a photo, showing it has no understanding of what is a body... (I can look those up if necessary).

And, of course, the answers are entirely untethered from reality - it is completely by chance whether the answer is correct or just wrong. It is run through a grammatical filter/generator at the end so it's usually grammatical, but no sort of truth filter (or ethical filter for that matter either).

I don't expect some abstract experience, I expect it to be able to break down it's work into fundamental abstract concepts and then construct an answer, and this it cannot do, or it would not be making these kinds of errors.

[0] https://twitter.com/Bestie_se_smeje/status/16210919157469184...

[1] https://twitter.com/albo34511866/status/1621608358003474432

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#278
post #262

Earlier quoted context omitted.

But that all boils down to are we having a scientific conversation or a philosophical conversation? In my opinion the only useful conversation is a scientific on. A philosophical conversation will and can never be resolved so if of no importance to this discussion. We can use philosophy to help guide our scientific conversation, but in the end only a scientific conversation can be helpful in reaching a meaningful/pra…

I think you're making a strawman to argue against. Nowhere above have I claimed that "knowing" requires "consciousness", or "it must be implemented identically to me to count", and in fact I believe neither. But: - In this context, following on the whole 2nd half of the 20th century where cognitive science and psychology moved past behaviorism and sought explanations of the _mechanisms_ underlying mental phenomena, a…

Yes valid points you make, but I feel they are still skipping something. To me it seems like you are asking "Does it know the same things we know?"> With the obvious answer is no because it doesn't have all of the senses we have.

Someone who is blind, doesn't have a lesser concept of knowing even though they are blind. They might not "know" things in the same way a someone who is seeing, but doesn't mean their version of knowing is any less, they just know fewer facts about the world. Specifically the visual facts of what things look like. Their "knowing" functionality is equal to someone who sees.

Similarly, someone who is blind, and deaf also has full ability for "knowing" even if they'll never know things in the visual or auditory spaces.

So my argument is that your premise is wrong, the fact that someone or something has fewer senses doesn't mean it's ability to know is any less.

So back to your LLM the fact it doesn't exists in the real world is not an exclusion from its ability to know. It does not need to have all of those experiences "to know". It will never know the physical meaning of concepts like we do. Just like I'll never know the details of a city block in Jakarta (as I've never been). But not having that experience (or any experiences of multiple senses) doesn't mean I don't know.

LLMs don't need multiple cross connected sensory experiences, nor extensive history with a physical or virtual world to know things.

For an entity "to know" it means it has a model it can use to make predictions.

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#279

Earlier quoted context omitted.

> Read to the end. The beginning is trivial the ending is unequivocal: chatGPT understands you. How does this necessarily and unequivocally follow from the blog post? All I see in it is a bunch of output formed by analogy: it has a general concept of what each command's output is kinda supposed to look like given the inputs (since it has a bajillion examples of each), and what an HTML or JSON document is kinda suppos…

Honestly I seriously find it hard to believe someone can read it to the end without mentioning how it queried itself. You're just naming the trivial things that it did. In the end It fully imagined a bash shell, an imaginary internet, an imaginary chatGPT on the imaginary internet, then on the imaginary chatGPT it created a new imaginary bash shell. The level of recursive depth here indicates deep understanding and s…

> In the end It fully imagined a bash shell, an imaginary internet, an imaginary chatGPT on the imaginary internet, then on the imaginary chatGPT it created a new imaginary bash shell.

In the general case, a shell is merely a particular prompt-response format with special verbs; the internet is merely a mapping from URLs to HTML and JSON documents; those document formats are merely particular facades for presenting information; and a "large language model" is merely something that answers free-form questions.

> The level of recursive depth here indicates deep understanding and situational awareness of what it is being asked. It demonstrates awareness of what "itself" is and what "itself" is capable of doing.

Uh, what? Why does that output require self-awareness? First, it's requested to produce the source of a document "https://chat.openai.com/chat". What might be behind such a URL? OpenAI Chat, presumably! And OpenAI is well known to create large language models, so a Chat feature is likely a large language model the user can chat with. Thus it invents "Assistant", and puts the description into the facade of a typical HTML document.

Then, it starts getting prompted with POST requests for the same URL, and it knows from the context of its previous output that the URL is associated with an OpenAI chatbot. So all that is left is to follow a regular question-answer format (since that's what large language models are supposed to do) and slap it into a JSON facade.

> But it MUST understand your query in order to produce the output show in the article. That much is obvious.

I'm saying that it "understands" your query only insofar as its words can be tied to the web of associations it's memorized. The impressive part (to me) is that some of its concepts can act as facades for other concepts: it can insert arbitrary information into an HTML document, a poem, a shell session, a five-paragraph essay, etc.

All of that can be achieved by knowing which concepts are directly associated with which other concepts, or patterns of writing. This is the reasoning by analogy that I refer to: if it knows what a poem about animals might look like, and it can imagine what kinds of qualities space ducks might possess, then it can transfer the pattern to create a poem about space ducks.

But none of this shows that it can relate ideas in ways more complex than the superficial, and follow the underlying patterns that don't immediately fall out from the syntax. For instance, it's probably been trained on millions of algebra problems, but in my experience it still tends to produce outputs that look vaguely plausible but are mathematically nonsensical. If it remembers a common method that looks kinda right, then it will always prefer that to an uncommon method.

I mean, it's not utterly impossible that GPT-4 comes along and humbles all the naysayers like myself with its frightening powers of intellect, but I won't be holding my breath just yet.

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#280

Earlier quoted context omitted.

People learn enormous amounts of things that we don’t actually “understand” in any deep way As long as our minds pops out appropriate thoughts for the given context we don’t even think about the magic machinery behind the scenes that did that. When queried about our thinking we are mostly creating a plausible story, not actually examining our own thinking. Also, blind people can talk sensibly about many visual phenom…

Blind people still have bodies and other sensory perceptions to relate visual meaning to. Temple Grandin is a high functioning autist who describes how visual thinkers translate words into pictures, because they think pictorially. LLMs don't have any embodied, grounded contact with the world, so their only understanding can be statistical/symbolic pattern matching of text. Which isn't how language works for humans, s…

Good points

But blind people can talk about color intelligently too, if not as completely as a sighted person. Despite not experiencing color qualia.

Post reply on HN