Theory of Mind May Have Spontaneously Emerged in Large Language Models
291–300 of 321 posts
Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models
#292Earlier quoted context omitted.
this is mind blowing to me. can anyone with more knowledge on the topic explain how ChatGPT is demonstrating this level of what seems like genuine understanding and reasoning? Like others I assumed that ChatGPT is gluing words together that commonly occur together. This is way more than that.
No, it's paraphrasing it's training data that likely contains these tasks in one form or another. Here's one I made : me : There's a case in the station and the policeman opens it near the fireman. The dog is worried about the case but the policeman isn't, what does the fireman think is in the station? chatgpt : As a language model, I do not have access to the thoughts of individuals, so I cannot say what the fireman…
Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models
#293Earlier quoted context omitted.
Defining what "knowing" is would be useful, yes, and analytic philosophers in epistemology do argue about this. One attribute that's classically part of the definition of "knowing" is that the thing which is known must be true. LLMs are pretty bad at this, but perhaps that can be fixed. But I would challenge you to imagine the situation the LLM is actually in. Do you understand Thai? If so, in the following, feel fre…
>But if we want AIs that 'know' the same things we know, then we have to build them to perceive in a multi-modal way, and interact with stuff in the world, rather than just self-supervising on piles of internet data. In other words, a LLM that is tied to a GAN that generates images, produces an system that can both describe to you what is a cat verbally and show you a picture of a cat. Does it, then, know what "a cat…
No, it's not the same LLM; you'd have to change the LLM in all of those cases. How does it receive input from the GAN? The typical LLM is constructed to literally receive a sequence of encoded tokens. There are vision transformers, and they do chunk images into tokens, and there are multimodal transformers, but none of these are fairly described as an LLM, and they're structurally different than something like ChatGPT. And after the structural changes, it would need to be trained on some new data that associates text sequences and image sequences, and after being optimized in that context you have a _different model_.
Does being able to identify images of cats mean the model knows what a cat is? No, and we could have said that a decade ago when deep learning for image classification was making its early first advances. Does being able to describe a cat from video mean you know what the cat is? Probably not, but maybe we're getting closer. Does knowing how to pet a cat mean you know what a cat is? Perhaps not if you need to be instructed to try to pet the cat.
But suppose 10 years from now, I have a domestic robot that has a vision system, and a motor control system, and an ability to plan actions and interact with a rich environment. I would say the following would be strong evidence of knowing what a cat is:
- it can not only identify or locate the cat, but can label parts of the cat, despite the cat having inconsistent shape. It can consistently pick up the cat in a way which is sensitive and considerate of the cat's anatomy (e.g. not by the head, by one paw, etc)
- it can entertain the cat, e.g. with a laser pointer, and can infer whether the cat is engaged, playful, stressed, angry etc
- it avoids placing fragile object near high edges, because it can anticipate that the cat is likely to knock them down, even if the cat is not currently near
- it can anticipate the cat's behavior and adjust plans around it; e.g. avoid vacuuming the sunny spot by the window in the afternoon when the cat is likely to be napping there
- it can anticipate the cat's reactions to stimuli, such as loud noises, a can of food opening, etc, and can incorporate these considerations into plans
Note, _none_ of the above have anything to do with language. If I add to the robot a bunch of NLP systems to hear and understand commands or describe its actions or perceptions, it may now know that a cat is called "cat", and how to talk about a cat, but these are distinct from knowing what a cat is.
Similarly,
- a human with some serious aphasia may be unable to describe the cat, but they can clearly still know what a cat is
- a dog can know what a cat is, in many important ways, despite having no language abilities
Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models
#294Earlier quoted context omitted.
I think you're making a strawman to argue against. Nowhere above have I claimed that "knowing" requires "consciousness", or "it must be implemented identically to me to count", and in fact I believe neither. But: - In this context, following on the whole 2nd half of the 20th century where cognitive science and psychology moved past behaviorism and sought explanations of the _mechanisms_ underlying mental phenomena, a…
Yes valid points you make, but I feel they are still skipping something. To me it seems like you are asking "Does it know the same things we know?"> With the obvious answer is no because it doesn't have all of the senses we have. Someone who is blind, doesn't have a lesser concept of knowing even though they are blind. They might not "know" things in the same way a someone who is seeing, but doesn't mean their versio…
> For an entity "to know" it means it has a model it can use to make predictions.
Great, every PID controller, every jupyter notebook or excel spreadsheet with a linear regression model, every count-down timer can make predictions and therefore "know" under this definition. But perhaps there's a broader class of things that "make predictions". Down this path lies panpsychism. When I throw a rock, its velocity in the x direction at time t is a great "predictor" of its velocity in the x direction at time t+delta, etc, etc. And maybe there's nothing inconsistent or fundamentally wrong with saying that every part of the physical universe "knows" at least something insofar as it participates in predicting or computing the future. But I think by so over-broadening the concept of knowing, it becomes useless, and impossible to make distinctions that matter.
Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models
#295Earlier quoted context omitted.
The problem with this facile view of things is that it seems to be a dead end for scientific theories. What if we just limited the science of birds to explaining how limb-flapping could produce levitation? Hmm yes. Birds are kind of like helicopters, it seems. Who’s to say that they are not basically one and the same? Moving on. If you are only interested in the most superficial tests and theories—like the Turing Tes…
> What if we just limited the science of birds to explaining how limb-flapping could produce levitation? Hmm yes. Birds are kind of like helicopters, it seems. Who’s to say that they are not basically one and the same? Moving on. Well said. I'm gonna steal this explanation. Also reminds me of the famous Carbonara quote: "if my grandmother had wheels, then she would be a bike" [1] [1] https://www.youtube.com/watch?v=A…
Well it could be argued that she would be a bike. Its possible to be multiple things at once. If she had 2 wheels and could be ridden by other humans to a destination she might qualify has a bike. She would also continue to be your grandmother.
Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models
#296Earlier quoted context omitted.
I wonder every time I see this take what it would mean under this definition of knowing things for a machine learning algorithm to ever know something. I find that especially important because to every appearance we are a machine learning algorithm. I don’t know how different the sort of knowing this algorithm has to the sort of knowing a human has, but you’re far more confident than I am that it’s a difference of ki…
The problem with this facile view of things is that it seems to be a dead end for scientific theories. What if we just limited the science of birds to explaining how limb-flapping could produce levitation? Hmm yes. Birds are kind of like helicopters, it seems. Who’s to say that they are not basically one and the same? Moving on. If you are only interested in the most superficial tests and theories—like the Turing Tes…
We could easily argue that birds are not a type of helicopter because for helicopter's we have a very specific set of flying properties required. It must have a main propeller for lift and a tail propeller to counter balance the main propeller from spinning the helicopter. If a bird flew with similar mechanism I would argue it was a helicopter.
We don't have a 100% accurate gauge for ToM as far as we know. This paper simply uses some of the best known tests for ToM and then states that either LLM can lead to emergent properties or that the current tests for ToM need to be re-thought.
Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models
#297This highlights one of the types of muddled thinking around LLMs. These tasks are used to test theory of mind because for people, language is a reliable representation of what type of thoughts are going on in the person's mind. In the case of an LLM the language generated doesn't have the same relationship to reality as it does for a person. What is being demonstrated in the article is that given billions of tokens o…
> These tasks are used to test theory of mind because for people, language is a reliable representation of what type of thoughts are going on in the person's mind. I actually think that human language is unreliable at expressing what's going on inside a persons mind[1]. My native language is not English, I have only introductory-level knowledge in the field of pragmatics[2], which makes me fully aware of the many way…
Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models
#298Earlier quoted context omitted.
>But if we want AIs that 'know' the same things we know, then we have to build them to perceive in a multi-modal way, and interact with stuff in the world, rather than just self-supervising on piles of internet data. In other words, a LLM that is tied to a GAN that generates images, produces an system that can both describe to you what is a cat verbally and show you a picture of a cat. Does it, then, know what "a cat…
At some point, when multiple components (including the LLM) have been connected to form a system that exhibits "knowing" (the way humans do), wouldn't the "intelligence" be distributed across the entire system rather than attributed primarily to the LLM? In other words, the LLM wouldn't be the equivalent of the human brain. Instead, it would just be equivalent to that part of the human brain that processes language.
Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models
#299Earlier quoted context omitted.
Yes valid points you make, but I feel they are still skipping something. To me it seems like you are asking "Does it know the same things we know?"> With the obvious answer is no because it doesn't have all of the senses we have. Someone who is blind, doesn't have a lesser concept of knowing even though they are blind. They might not "know" things in the same way a someone who is seeing, but doesn't mean their versio…
I think your argument goes off the rails when it jumps from "you don't need any particular sense modality to know" to "you don't need any percepts, or experience of reality or simulated unreality to know". That's a big leap, and I can't disagree more. > For an entity "to know" it means it has a model it can use to make predictions. Great, every PID controller, every jupyter notebook or excel spreadsheet with a linear…
I still feel this the the point where you're making a difference based on you desired outcome vs the actual system. ChatGPT absolutely does have precepts / a sense. It has a sense of "textual language". It also has a level of sequencing or time w.r.t. word order of that text.
While you're saying experience, it seems like in your definition experience only counts if there is a spatial component to it. Any experience without a physical spatial component to you seems like it's not valid sense or perception.
Again taking this in the specific, imagine someone could only hear via one ear, and that is their only sense. So there is no multi-dimensional positioning of audio, just auditory input. It's clear to me that person can still know things. Now if you also made all audio the same loudness so there is no concept of distance with it, it still would know things. This is now the same a simple audio stream, just like ChatGPT's langauge stream. Spatial existence is not required for knowledge. And from what I'm understanding that is what underpins your definition of a reality/experience (whether physical or virtual).
Or as a final example lets say you are Magnus Carlson. You know a ton about chess, best in the world. You know so much about chess that you can play entire games via chess notation (1. e4, e6 2. d4 e5 ...). So now an alternate world where there is even a version of Magnus that has never sat in front of a chess board and only ever learned chess by people reciting move notation to him. Does the fact that no physical chess boards exist and there is no reality/environment where chess exists mean he doesn't know chess? Even if chess were nothing but streams of move notations it still would be the same game, and someone could still be an expert at it knowing more than anyone else.
I feel your intuition is leading your logic astray here. There is no need for a physical or virtual environment/reality for something to know.
Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models
#300Earlier quoted context omitted.
Honestly I seriously find it hard to believe someone can read it to the end without mentioning how it queried itself. You're just naming the trivial things that it did. In the end It fully imagined a bash shell, an imaginary internet, an imaginary chatGPT on the imaginary internet, then on the imaginary chatGPT it created a new imaginary bash shell. The level of recursive depth here indicates deep understanding and s…
> In the end It fully imagined a bash shell, an imaginary internet, an imaginary chatGPT on the imaginary internet, then on the imaginary chatGPT it created a new imaginary bash shell. In the general case, a shell is merely a particular prompt-response format with special verbs; the internet is merely a mapping from URLs to HTML and JSON documents; those document formats are merely particular facades for presenting i…
Your attempt to trivialize it doesn't make any sense. It's like watching someone try to trivialize the moon landing. "Oh all we did was put a bunch of people in some metal cylinder then light the tail end on fire. Boom simple propulsion! and then we're off to the moon! You don't need any intelligence to do that!"
>I'm saying that it "understands" your query only insofar as its words can be tied to the web of associations it's memorized. The impressive part (to me) is that some of its concepts can act as facades for other concepts: it can insert arbitrary information into an HTML document, a poem, a shell session, a five-paragraph essay, etc.
You realize the human brain CAN only be the sum of it's own knowledge. That means anything creative we produce anything at all that comes from the human brain is DONE by associating different things together. Even the concept of understanding MUST be done this way simply because the human brain can only create thoughts by transforming it's own knowledge.
YOU yourself are a web of associations. That's all you are. That's all I am. The difference is we have different types of associations we can use. We have context of a three dimensional world with sound, sight and emotion. chatGPT must do all of the same thing with only textual knowledge and a more simple neural network so it's more limited. But the concept is the same. YOU "understand" things through "association" also because there is simply no other way to "understand" anything.
If this is what you mean by "reasoning by analogy" then I hate to tell you this, but "reasoning by analogy" is "reasoning" in itself. There's really no form of reasoning beyond associating things you already know. Think about it.
>But none of this shows that it can relate ideas in ways more complex than the superficial, and follow the underlying patterns that don't immediately fall out from the syntax. For instance, it's probably been trained on millions of algebra problems, but in my experience it still tends to produce outputs that look vaguely plausible but are mathematically nonsensical. If it remembers a common method that looks kinda right, then it will always prefer that to an uncommon method.
See here's the thing. Some stupid math problem it got wrong doesn't change the fact that the feat performed in this article is ALREADY more challenging then MANY math problems. You're dismissing all the problems it got right.
The other thing is, I feel it knows math as well as some D student in highschool. Are you saying the D student in highschool can't understand anything? No. So you really can't use this logic to dismiss LLMs because PLENTY of people don't know math well either, and you'd have to dismiss them as sentient beings if you followed your own reasoning to the logical conclusion.
>I mean, it's not utterly impossible that GPT-4 comes along and humbles all the naysayers like myself with its frightening powers of intellect, but I won't be holding my breath just yet.
What's impossible here is to flip your bias. You and others like you will still be naysaying LLMs even after they take your job. Like software bugs, these AIs will always have some flaws or weaknesses along some dimension of it's intelligence and your bias will lead you to magnify that weakness (like how you're currently magnifying chatGPT's weakness in math). Then you'll completely dismiss the fact that chatGPT taking over your job as some trivial "word association" phenomenon. There's no need to hold your breath when you wield control of your own perception of reality and perceive only what you want to perceive.
Literally any feat of human intelligence or artificial intelligence can literally be turned into a "word association" phenomenon using the same game you're running here.