Live data from Hacker News

Large language models lack deep insights or a theory of mind

arxiv.org

141–150 of 270 posts

Re: Large language models lack deep insights or a theory of mind

#141
post #86

Here's my theory: Consider a typical LLM token vector used to train and interact with an LLM. Now imagine that other aspects of being human (sensory input, emotional input, physical body sensation, gut feelings, etc.) could be added as metadata to the the token stream, along with some kind of attention function that amplified or diminished the importance of those at any given time period -- all still represented as a…

So my observation is that we could embody an AI so that it learns theory of mind-body--but then we could remove the body. This gives a theory of mindful entity that does not need a body to exist.

Then the next research step could be to study those properties so as to reconstruct/reproduce a theory of mind-body AI, without needing any embodiment process at all to obtain it. Is that, in principle, possible? It is unclear me.

Re: Large language models lack deep insights or a theory of mind

#142
post #9

I appreciate this paper for relatively clearly stating what "human-like" might entail, which in this case involves "reasoning about the causes behind other people's behavior" which is "critical to navigate the social world" as outlined in this citation: https://www.sciencedirect.com/science/article/abs/pii/S00100... I get frustrated often when people argue "well, it isn't really intelligent" and then give examples th…

The underlying problem is that "intelligence" is itself a crappy, poorly defined word with a fraught and inconsistent history. It doesn't appear until the early 20th century, in the shadow of compulsory education and the challenges it presented, first as a technical label for attempts to sort students -- and later soldiers -- into the tracks in which they're most likely to succeed, and then being haphazardly asserted…

The way you're using it when you worry about "super-intelligence" is in the sense of intelligence being some universal, unbounded, quantitative independent variable along the lines of "the more intelligent something is, the more cunningly it can pursue some rationalized goal" -- some master strategist.

I appreciate the highlighting of the term intelligence being ill-defined. Moreover, it's certainly true that "AI safety analysts" takes intelligence as a sort magic wand term and this seems to drive their arguments.

All that said, since both computers and human brains are material artifacts, it doesn't seem impossible to create a device that combines their properties. It seems plausible that such a thing could have a variety of dangers.

For many, an "super-intelligent" software whose "motives" we don't understand is just a program that produces incorrect outputs and ought to be debugged or retired, and the more interesting questions around machine "intelligence" are practical ones like "what tasks are these programs well-suited for".

We saw early Bing Chat behave, not in ways we couldn't understand but like a deranged and vengeful human. Certainly, it was merely simulating human behavior but if today's methods produce artifacts that unselectively amplify human behaviors, it's not hard to imagine problems appearing.

We can hope that there's a fundamental difference between programs that simulate human language and programs able to plan and carry out long term goals (and carrying out long term goals is something people do so there's no good reason some kind of program couldn't do that).

I think you're right that particular weirdness of the "doomers" makes some other portion of the population dismiss concerns. But that isn't an argument that the doom isn't possible - it should be an argument to clarify how we talk of computation and human capacities (see, I don't to say "intelligence" unless I want to).

Re: Large language models lack deep insights or a theory of mind

#143
post #37

Earlier quoted context omitted.

> Maybe the soul is social My pet theory about human consciousness is that is that consciousness is simply recursive theory of mind. Theory of mind [1] is our ability to simulate and reason about the mental states of others. It's how we predict what people are thinking and how they will react to our actions, which is critical for choosing how to act in a social environment. But when you're thinking about what's in so…

Theory of mind is interesting but one wouldn't want to hinge consciousness upon it. That direction would likely contain weird outcomes if the science progressed, something like "Dogs are barely-conscious due to their pack structure, they have a couple levels of recursive theory of mind but they can't sustain it as deep as we can. But cats didn't have that pack structure, they're not conscious at all." Or, "this perso…

"Consciousness" is one of those loaded words that means a few different things in different contexts. When I say consciousness is about recursive theory of mind, I'm not trying to say that when you're asleep you're unable to do social reasoning. That's a different use of the same word.

I mean "conscious" in the sense of self-awareness or sentience.

I'm also not ascribing any moral or cognitive superiority or inferiority to different levels of it. The fact that a cat might be less self-aware of its suffering because it think about how it would explain its pain to other cats does mean imply that I'm saying it should be OK to torture cats.

I'm just interested in what is going in human brains when we "feel self-aware". What are we doing when we're thinking about what we're doing? Where does that sense of perceptual distance come from when we are aware of ourselves? And my pet theory is that the distance comes from imagining how we look through others' eyes and developed from our highly advanced social reasoning.

Re: Large language models lack deep insights or a theory of mind

#144
post #93

Earlier quoted context omitted.

Couldn't agree more. How about this -- I think we've already reached AGI. Let me know if this tracks: Pick a set of tasks that can be considered AGI tasks. Provided the task sequences can be compared as closer to AGI or further from AGI, we can create a reward model using the same techniques as were used by ChatGPT via RLHF. Thus, for any definition of AGI that is meaningful and selectable, even if subjectively selec…

Well... humans have different mental "spaces" (not intended as a technical term). Let's say I'm deep in a coding problem. A co-worker comes by and says "How did your team do in the game yesterday?". I say, "Um, uh... sorry, my head's not there right now." It takes us time to swap between mental "spaces". So, if I have an AGI (defined as having a trained model for almost everything, even if that turns out to be a larg…

The power of analogy is one of the most important things that humans seem to have.

Humans typically use the toolset they've seen along the way to solve problems (hence if you have a hammer all problems become nails statement). When you get people that are multi-disciplinary they commonly can solve a complex problem in one field by bringing parts of solutions from other fields.

Hence if you have more life experiences (especially positive/learning ones) you are typically better off then a person who does not.

Also I think this is where a lot of interest in Q* learning after the OpenAI thing occurred, as this would be a means of allowing an AI to explore problem spaces and enlist specialist AI and tools for it to do so.

Re: Large language models lack deep insights or a theory of mind

#145

Earlier quoted context omitted.

> Maybe the soul is social My pet theory about human consciousness is that is that consciousness is simply recursive theory of mind. Theory of mind [1] is our ability to simulate and reason about the mental states of others. It's how we predict what people are thinking and how they will react to our actions, which is critical for choosing how to act in a social environment. But when you're thinking about what's in so…

You might find this book interesting! This is essentially the theory put forward. https://www.google.com/books/edition/Consciousness_and_the_S...

Ah, that looks perfect! Thank you! I knew other people smarter than me must have stumbled onto this idea as well.

Re: Large language models lack deep insights or a theory of mind

#146

For me, the entire AGI conversation is hyperbolic / hype. How can we infer intelligence to something when we, ourselves, have such a poor (none) grasp of what makes us conscience? I'm associating intelligence with consciousness - because it seems correlated. Are we really ready to associate "AGI" with solving math problems ("new Q algo.")? That seems incredibly naive & reinforces my opinion that LLM's are much more l…

It's not hype. It's a language problem that makes people like you think this way.

The problem is consciousness is a vocabulary word that establishes a hard boundary where such a boundary doesn't exit. The language makes you think either something is conscious or it is not when the reality is that these two concepts are actually extreme endpoints on a gradient.

The vocabulary makes the concept seem binary and makes it seem more profound then it actually is.

Thus we have no problem identifying things at the extreme. A rock is not conscious. That's obvious. A human IS conscious, that's also obvious. But only because these two objects are defined at the extremes of this gradient.

For something fuzzy like chatGPT, we get confused. We think the problem is profound, but in actuality it's just poorly defined vocabulary. The word consciousness, again, assumes the world is binary that something is either/or, but, again, the reality is a gradient.

When we have debates about whether something is "conscious" or not we are just arguing about where the line of demarcation is drawn along the gradient. Does it need a body to be conscious? Does it need to be able to do math? Where you draw this line is just a definition of vocabulary. So arguments about whether LLMs are conscious are arguments about vocabulary.

We as humans are biased and we blindly allow the vocabulary to mold our thinking. Is chatGPT conscious? It's a loaded question based on a world view manipulated by the vocabulary. It doesn't even matter. That boundary is fuzzy, and any vocab attempting to describe this gradient is just arbitrary.

But hear me out. chatGPT and DALL-E is NOT hype. Why? Because along that gradient it's leaps and bounds further than anything we had just even a decade ago. It's the closest we ever been to the extreme endpoint. Whichever side you are on in the great debate both sides can very much agree with this logic.

Re: Large language models lack deep insights or a theory of mind

#147

Earlier quoted context omitted.

> Maybe the soul is social My pet theory about human consciousness is that is that consciousness is simply recursive theory of mind. Theory of mind [1] is our ability to simulate and reason about the mental states of others. It's how we predict what people are thinking and how they will react to our actions, which is critical for choosing how to act in a social environment. But when you're thinking about what's in so…

If this is correct, do you think GPT5 will be conscious because its training data will include a lot of itself (albeit GPT4 not 5)

I'm sorry, but I honestly don't find philosophical questions about the intelligence or sentience of generative AI interesting at all.

Re: Large language models lack deep insights or a theory of mind

#148
post #133
post #114

Earlier quoted context omitted.

Completely agree. From a computer science point of view: a single prompt/response cycle from a LLM is equivalent to a pure function; the answer is a function of the prompt and the model weights and is fundamentally reducible to solving a big math equation (in which each model parameter is a term.) It seems almost self evident that "reasoning" worthy of the name would involve some sort of iterative/recursive search pr…

Ideally a recursive execution would also be a pure function - maybe a better way to put it about current LLMs is that they are a single mathematical expression being built up from a fix number of nodes and only addition and multiplication.

yes, except the "reasoning" process should also be able to look up facts (retrieval) and invoke external tools, making it non-pure.

Re: Large language models lack deep insights or a theory of mind

#149

Earlier quoted context omitted.

> It's not. They don't realize it, they're merely referring to stopping your internal monologue. They certainly have realized that. It's one of the first things you notice doing awareness meditation; thoughts appear from nowhere even if you didn't try to think them.

[flagged]

That seems like an unusually rude thing to say about Buddha. You might be experiencing dukkha.

Re: Large language models lack deep insights or a theory of mind

#150
post #136

Earlier quoted context omitted.

Or there are enough of those examples in the training set that it can guess well. Not sure how such an example would prove anything when we know an LLM is just guessing the best words. Nothing I’ve seen shows evidence of any sort of abstract concepts in there.

Wouldn't this also be the same for humans?

If you introspect and decide it is so, I won't disagree with you.
Post reply on HN