Live data from Hacker News

Theory of Mind May Have Spontaneously Emerged in Large Language Models

arxiv.org

251–260 of 321 posts

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#251

Earlier quoted context omitted.

Like GP said, the LLM has no chance at knowing what a cat is, regardless of how much data it ingests, because a cat is not made of data. It's not like you're getting closer and closer to knowing what a "Mæw" is. You were at the same remote distance all the time. This is called the "grounding problem" in AI. As for how you would test it, I think one-shot learning would get one closer to proving understanding.

The grounding problem is an intelligence problem, not an artificial intelligence problem. How would you envision a test based on one-shot learning working?

The question of grounding is a problem that arises in thinking about cognition in general, yes. In AI, it changes from a theoretical problem to a practical one, as this whole discussion proves.

As for one-shot learning, what I was driving at, is that a truly intelligent system should not need to consume millions of documents in order to predict that, say, driving at night puts larger demands on one's vision than driving during the day. Or any other common sense fact. These systems require ingesting the whole frickin' internet in order to maybe kinda sometimes correctly answer some simple questions. Even for questions restricted to the narrow range where the system is indeed grounded: the world of symbols and grammar.

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#252

Earlier quoted context omitted.

Generally I don't buy these arguments which require embodiment, because they don't seem to align well to what else I know about my world. Rather than your Thai text example, let's consider a friend of my sister H. H has been profoundly blind from birth. Not "legally blind" with the world a blur, her eyes actually don't work. Direct lived experience of a summer day is to her literally just feeling warmth on her face f…

That's an argument from ignorance, and it's not credible. The potential total scope of experience is irrelevant. The reality is that you have an embodied experience of purple shared with most humans. Unfortunately your sister doesn't. She will have a linguistic placeholder for the concept of purple, probably surrounded by verbal associations. But that's all. It's an ironically apt analogy, because ChatGPT has the lin…

There's a few particular problems we have with the word intelligence/sentience, mostly revolving around that we evolved embodiment first and then added more and more complex intelligence/sentience on top of an ever changing DNA structure.

Much like when humans started experimenting with flight we tried to make flapping things like birds, but in the end it turns out spinning blades gives us capabilities above and beyond bodies that flap.

Back to the embodiment problem. For us as humans we have limits like only having one body. It has a great number of sensors but they are still very limited in relation to what reality has to offer, hence we extend our senses with technology. And with that there is no reason machine intelligence embodiment has to look anything like ours. Machine intelligence could have trillions of sensors spread across the planet as an example.

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#253
post #204

Earlier quoted context omitted.

Seems like that's a consequence of the philosophical semantics of the word "know", not really a statement about the demonstrable capabilities of the LLM. In other words, why does it matter?

In the context of a discussion on whether LLMs could have a theory of mind? I think the ability to know anything at all matters to evaluate that conclusion. More generally, what an LLM actually knows or understands is important if you're considering using one for anything other than generating first drafts which will be fact checked by humans.

If you're depending on fact checking by any one human I think that the last few years in politics should be a sufficient warning to the dangers of that. In the end the LLM will have to be integrated into larger systems that cross check each other.

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#254

Earlier quoted context omitted.

I wonder every time I see this take what it would mean under this definition of knowing things for a machine learning algorithm to ever know something. I find that especially important because to every appearance we are a machine learning algorithm. I don’t know how different the sort of knowing this algorithm has to the sort of knowing a human has, but you’re far more confident than I am that it’s a difference of ki…

I am not in the field so I cannot speak very eloquently what it would mean for machine learning algorithm to "ever know something". But I feel that e.g. Simulations and perhaps expert systems of yore, were qualitatively closer to getting there. Their error modes were radically different. They started inductively with rules, rather than arriving at them statistically almost by accident.

Did we end up with human intelligence statistically pretty much by accident?

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#255

Earlier quoted context omitted.

We cannot assume that, because text generation is all these models do, then it must be possible to get answers to the questions we want to ask by examining their textual responses. It is fair to ask why, if we accept these verbal challenges as good evidence for a theory of mind in children, we would not accept them for these models, but children have nothing like the memory for text that these models have, and the co…

> To be clear, I am not arguing that it would be impossible to show a theory of mind in a system that can only interact through text I think you are, because > a model with greater capabilities than responding to prompts interacts in other ways than text. Even then, I don't see what's so special about language that it needs to be separated from other ways of interaction. If language is not enough to derive empirical…

Two models having a coherent conversation - a scenario which follows directly from my post - would be a purely textual example of what I mean.

> Perhaps we will never find out if language models have a theory of mind.

We appear to be in agreement here.

When the state of our knowledge is 'maybe', it seems rash to assume either 'yes' or 'no'.

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#256
post #141

Earlier quoted context omitted.

Maybe the participant "knows" something about the Thai language? But that's different from knowing anything about the things being discussed. The jumping off point for this, which motivated a question about what it is to know, was the comment: > What this article is not showing (but either irresponsibly or naively suggests) is that the LLM knows what a bag is, what a person is, what popcorn and chocolate are, and can…

Language doesn’t come first for humans. Experiencing the world does. Languages then become symbols to communicate experiencing the world through our senses and emotional/mental states. I’m not sure why people get hung up on language models not being the same thing when they start and end with language.

Indeed. Give a model some kind of autonomous sensors, make it stateful with memory and continuous retraining, make it possible for it to act and learn from its actions, maybe even model some kind of hormonal influence etc. and I'm pretty sure that at some point an actual Theory of Mind will actually emerge and we'll be debating what kind of legal rights such a model should possess. We're pretty clearly not at that point yet.

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#257

Earlier quoted context omitted.

The grounding problem is an intelligence problem, not an artificial intelligence problem. How would you envision a test based on one-shot learning working?

The question of grounding is a problem that arises in thinking about cognition in general, yes. In AI, it changes from a theoretical problem to a practical one, as this whole discussion proves. As for one-shot learning, what I was driving at, is that a truly intelligent system should not need to consume millions of documents in order to predict that, say, driving at night puts larger demands on one's vision than driv…

Why do you believe that a system should not need to consume millions of documents in order to be able to make predictions?

For your example, the concepts of driving, night, vision, all need to be clearly understood, as well as how they relate to each other. The idea of 'common sense' is a good example of something which takes years to develop in humans, and develops to varying extents (although driving at night vs at day is one example, driving while drunk and driving while sober is a different one where humans routinely make poor decisions, or have incorrect beliefs).

It's estimated that humans are exposed to around 11 million bits of information per second.

Assuming humans do not process any data while they sleep (which is almost certainly false): newborns are awake for 8 hours per day, so they 'consume' around 40GB of data per day. This ramps up to around 60GB by the time they're 6 months old. That means that in the first month alone, a newborn has processed 1TB of input.

By the age of six months, they're between 6 and 10TB, and they haven't even said their first word yet. Most babies have experienced more than 20TB of sensory input by the time they say their first word.

Often, children are unable to reason even at a very basic level until they have been exposed to more than 100TB of sensory input. GPT-3, by contrast was trained on a corpus of around 570GB worth of text.

We are simply orders of magnitude away from being able to make a meaningful comparison between GPT-3 and humans and determine conclusively that our 'intelligence' is of a different category to the 'intelligence' displayed by GPT-3.

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#258
post #33

This highlights one of the types of muddled thinking around LLMs. These tasks are used to test theory of mind because for people, language is a reliable representation of what type of thoughts are going on in the person's mind. In the case of an LLM the language generated doesn't have the same relationship to reality as it does for a person. What is being demonstrated in the article is that given billions of tokens o…

> What this article is not showing (but either irresponsibly or naively suggests) is that the LLM knows what a bag is, what a person is, what popcorn and chocolate are, and can then put itself in the shoes of someone experiencing this situation, and finally communicate its own theory of what is going on in that person's mind. That is just not in evidence.

You can't really conclude that unless you think we have a deep mechanistic understanding of "knowing". I agree that LLM doesn't have the same knowledge of these things as a human does, but it clearly has some kind of knowledge of how these words relate to each other. It "knows" that a "person" "puts" "things" "in" "bags", and for instance, that bags don't put things in people. So it clearly has some knowledge of bags and people, it just doesn't have multisensory associations with these objects.

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#259

Earlier quoted context omitted.

> To be clear, I am not arguing that it would be impossible to show a theory of mind in a system that can only interact through text I think you are, because > a model with greater capabilities than responding to prompts interacts in other ways than text. Even then, I don't see what's so special about language that it needs to be separated from other ways of interaction. If language is not enough to derive empirical…

Two models having a coherent conversation - a scenario which follows directly from my post - would be a purely textual example of what I mean. > Perhaps we will never find out if language models have a theory of mind. We appear to be in agreement here. When the state of our knowledge is 'maybe', it seems rash to assume either 'yes' or 'no'.

What does it change when you add another model? I don't see how this lets us extract extra information.

What distinguishes two conjoined models from one model with a narrowing across the middle?

If the idea is to have two similar minds building a theory of each other, then I guess this could be informative, but first we'd have to establish that the models are "minds" in the first place. It's not clear to me what that requires.

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#260

Earlier quoted context omitted.

May I play devil's advocate? The fallacy of this paper granted, is it worth questioning our belief that there is more to intelligence than the "appearance of intelligence"? What if the lack of hallucination in human being is due to our self-imposed guard (hello, frontal cortex) that is developed via an evolutionary process (aka, biological reinforcement training)? To stretch the argument a bit further, what if halluc…

To be honest, I think your question is the other side of the conceptual coin here -- either the article is wrong and GPT3 isn't "sentient," or the article is right and we need to radically upend our concept of sentience, probably via an eliminativist materialism that ends in at least epiphenomenalism if not full-bore illusionism, removing either free will or consciousness itself from our worldview. To be honest, I fi…

> To be honest, I think your question is the other side of the conceptual coin here -- either the article is wrong and GPT3 isn't "sentient," or the article is right and we need to radically upend our concept of sentience

We don't really have a mechanistic understanding of sentience, so I'm not sure there's much to overturn. This is why I'm so annoyed every time some "expert" claims that LaMDA or GPT are not sentient or not conscious or what have you. These are all vague concepts that lack mechanistic definitions that would let us make such definitive claims.

That said, consciousness is almost certainly an illusion, and it will be a particular kind of information processing system with certain properties [1]. GPT and other LLMs may or may not qualify, time will tell.

[1] https://www.pnas.org/doi/10.1073/pnas.2116933119

Post reply on HN