Live data from Hacker News

Theory of Mind May Have Spontaneously Emerged in Large Language Models

arxiv.org

301–310 of 321 posts

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#301

Earlier quoted context omitted.

A child can 'see' maths though, they can see that if you have one apple over here and one orange over there, then you have two pieces of fruit all together. If you only ever allowed a child to read about adding, without ever being able to physically experiment with putting pieces together and counting them, likely children would not be able to add either. In fact, many teachers and schools teach children to add using…

Yet ChatGPT totally - apparently - gets 1 + 1. In fact it aces the addition table way beyond what a child or even your average adult can handle. It's only when you get to numbers in the billions that it's weaknesses become apparent. One thing it starts messing up is carry-over operations, from what I can see. Btw. the treshold used to be significanly lower, yet that doesn't convince me in the least that it's made pro…

That’s an interesting read, thank you. But my question is a bit more fundamental than that.

Ultimately, my point is that although the argument is that an LLM doesn’t “know” anything, I am not sure that there is something categorically different in terms of what we “know” vs what an LLM “knows”, we have just had more training on more different types of data (and the ability to experiment for ourselves).

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#302
Clarification: An LLM doesn't have a 'Theory of Mind', it just looks like one. Maybe you're thinking of the Chinese room analogy. But this isn't about the Chinese room, it's about "measuring any metric is only effective until you optimize for that metric" problem.

Analogy: An autistic person of normal intelligence who is obsessed with problems and solutions for ToM may be good at solving them but still not have ToM.

Do I understand well?

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#303

Earlier quoted context omitted.

Honestly I seriously find it hard to believe someone can read it to the end without mentioning how it queried itself. You're just naming the trivial things that it did. In the end It fully imagined a bash shell, an imaginary internet, an imaginary chatGPT on the imaginary internet, then on the imaginary chatGPT it created a new imaginary bash shell. The level of recursive depth here indicates deep understanding and s…

> In the end It fully imagined a bash shell, an imaginary internet, an imaginary chatGPT on the imaginary internet, then on the imaginary chatGPT it created a new imaginary bash shell. In the general case, a shell is merely a particular prompt-response format with special verbs; the internet is merely a mapping from URLs to HTML and JSON documents; those document formats are merely particular facades for presenting i…

Another link for you:

https://news.ycombinator.com/news

Llms (the exact same architecture as chatGPT) trained to use calculators. Tell me which one requires "understanding". Learning how to use a calculator or learning how to do math perfectly?

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#304
post #33

This highlights one of the types of muddled thinking around LLMs. These tasks are used to test theory of mind because for people, language is a reliable representation of what type of thoughts are going on in the person's mind. In the case of an LLM the language generated doesn't have the same relationship to reality as it does for a person. What is being demonstrated in the article is that given billions of tokens o…

Today at the zoo I saw chimpanzees and right next to their area there was a fun fact tablet. It said that before around 1960 it was thought that humans were the only species to use tools. After discovering the same for chimps, Louis Leaky said “Now we must redefine tool, redefine man, or accept chimpanzees as human.”

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#305

Earlier quoted context omitted.

But that all boils down to are we having a scientific conversation or a philosophical conversation? In my opinion the only useful conversation is a scientific on. A philosophical conversation will and can never be resolved so if of no importance to this discussion. We can use philosophy to help guide our scientific conversation, but in the end only a scientific conversation can be helpful in reaching a meaningful/pra…

> Something that is able to simulate having a theory of mind sufficiently well does actually have a theory of mind. That presupposes that our existing tools for detecting the presence of ToM are 100% accurate. Might it be possible that they are imprecise and it’s only now that their critical flaws have been exposed?

But if our understanding of ToM is so flawed in practice, what does it say about all the confident proclamations that AIs "aren't real" because they don't have it?

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#306

Earlier quoted context omitted.

Yeah but if you ask the model what a cat is, it'll use other words that describe a cat because they're usually used in a sentence about cats. These words must relate to cats. So if I ask you what a cat is, you'll use words that relate to cats. Sure, you may visually see these words in your head. You may visually see a cat in your head, but your output to me is just a description of a cat. That's the same thing the ne…

But for us a cat is a living creature we interact with, not simply a description. We understand people's reactions to cats based on human-animal interactions, particularly as cute pets, not because of language prediction of what a cat description would be. People usually have feelings about cats, they have conscious experiences of cats, they often have emotional bonds with cats (or dislike them), they may be allergic…

Not "for us"; only for those of us who have, in fact, been exposed to cats.

And why do you think "feeling of a cat" cannot be encoded as a stream of tokens?

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#307

Earlier quoted context omitted.

A follow up question: as a human doesn't start with "knowing" something either and first creates definitions for objects or words, which it then uses to build increasingly abstract concepts that we eventually classify as "knowledge" on the thing, is there anything that would stop LLMs from being able to do the same thing? I fully agree the capability is not there yet, but I can't say what would stop an appropriately…

Socrates argued that we are born knowing everything, but we forgot most if it. Learning is simply the act of recalling what you once knew. The point, for this thread, is not whether or not Socrates was correct. Rather, it’s a warning that we must not confidently assume we are anything like a machine. We may have souls, we may be eternal, there may be something utterly immaterial at the heart of us. As we strive to un…

And, conversely, we might just be so full of ourselves that we are willing resort to claims on the immaterial if that's what it takes to not give up the exceptionalism.

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#308

My belief, based on experiences with domestic and wild animals is that there is nothing uniquely human about "theory of mind". It's a running gag in our household (where my wife runs a riding academy) that academics just published a paper showing that some animal (e.g. horse) has just been proven to have some cognitive capability that seems pretty obvious if you work with those animals. It's very hard to know what is…

I just had a chat with my kindergartner about humans being a kind of mammal, and out of some dark, musty recess of my brain crawled the sound of "You are a human animal"[1] You are a human animal You are a very special breed For you are the only animal Who can think, Who can reason, Who can read. Now all your pets are smart, that's true! But none of them can add up 2 and 2 Because the only thinking animal Is You! You…

https://en.wikipedia.org/wiki/Ren%C3%A9_Descartes#On_animals

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#309
post #65

Earlier quoted context omitted.

Here's a simpler scenario that doesn't involve reading: Me: There is a banana on the table. Someone comes and peels the banana and shows you that inside it, there is actually a carrot! Then they carefully stick the peel back so it look unpeeled. What is inside the banana skin? ChatGPT: According to the scenario described, there is a carrot inside the banana peel that has been carefully placed back to look unpeeled. M…

this is mind blowing to me. can anyone with more knowledge on the topic explain how ChatGPT is demonstrating this level of what seems like genuine understanding and reasoning? Like others I assumed that ChatGPT is gluing words together that commonly occur together. This is way more than that.

It's not an either-or.

What we're doing with LLMs is, in some sense, an experiment in extremely lossy compression of text. But what if the only way you can compress all those hundreds of terabytes of text is by creating a model of the concepts described by that text?

Re: Theory of Mind May Have Spontaneously Emerged in Large Language Models

#310
post #294

Earlier quoted context omitted.

I think your argument goes off the rails when it jumps from "you don't need any particular sense modality to know" to "you don't need any percepts, or experience of reality or simulated unreality to know". That's a big leap, and I can't disagree more. > For an entity "to know" it means it has a model it can use to make predictions. Great, every PID controller, every jupyter notebook or excel spreadsheet with a linear…

> you don't need any percepts, or experience of reality or simulated unreality to know". That's a big leap, and I can't disagree more. I still feel this the the point where you're making a difference based on you desired outcome vs the actual system. ChatGPT absolutely does have precepts / a sense. It has a sense of "textual language". It also has a level of sequencing or time w.r.t. word order of that text. While yo…

You're still fighting a strawman. You're the only participant in this thread that's talking about space. I'm going to discontinue this conversation with this message since (aptly), you seem happy responding to views whether or not they come from an actual interlocutor.

- I disagree that inputs to an LLM as a sequence of encoded tokens constitute a "a sense" or "percepts". If inputs are not related to any external reality, I don't consider those to be perception, any more than any numpy array I feed to any function is a "percept".

- I think you're begging the question by trying to start with a person and strip down their perceptual universe. I think that comes with a bunch of unstated structural assumptions which just aren't true for LLMs. I think space/distance/directionality aren't necessary for knowing some things (but bags, chocolate and popcorn as lsy raised at the root of this tree probably require notions of space). I can imagine a knowing agent whose senses are temperature and chemosensors, and whose action space is related to manipulating chemical reactions, perhaps. But I think action, causality and time are important for knowing almost anything related to agenthood, and these are structurally absent in ChatGPT UUIC. The RLHF loop used for Instruct/ChatGPT is a bandit setup. The "episodes" it's playing over are just single prompt-response opportunities. It is _not_ considering "If I say X, the human is likely to respond Y, an I can then say Z for a high reward". Though we interact with ChatGPT through a sequence of messages, it doesn't even know what it just said; my understanding is the system has to re-feed the preceding conversation as part of the prompt. In part, this is architecturally handy, in that every request can be answered by whichever instance the load-balancer picks. You're likely not talking to the same instance, so it's good that it doesn't have to reason about or model state.

But I actually think both of these are avenues towards agents which might actually have a kind of ToM. If you bundled the transformer model inside a kind of RNN, where it could preserve hidden state across the sequence of a conversation, and if you trained the RLHF on long conversations of the right sort, it would be pushed to develop some model of the person it's talking to, and the causes between its responses and the human responses. It still wouldn't know what a bag is, but it could better know what conversation is.

Post reply on HN