Live data from Hacker News

Large language models lack deep insights or a theory of mind

arxiv.org

81–90 of 270 posts

Re: Large language models lack deep insights or a theory of mind

#81
post #63

Earlier quoted context omitted.

The underlying problem is that "intelligence" is itself a crappy, poorly defined word with a fraught and inconsistent history. It doesn't appear until the early 20th century, in the shadow of compulsory education and the challenges it presented, first as a technical label for attempts to sort students -- and later soldiers -- into the tracks in which they're most likely to succeed, and then being haphazardly asserted…

> It [the word "intelligence"] doesn't appear until the early 20th century I'm not sure what you mean here, since the word dates back to the late 14th century with roughly the same meaning as now. Perhaps you're thinking of "intelligence quotient"? https://www.etymonline.com/word/intelligence

Summary etymology can provide interesting reference points when looking a the history of ideas, but isn't sufficient because adjacent concepts change their meaning over time as well. It's good for showing when a word was attested and where to start looking for an understanding of how it was used and considered.

Where you say "roughly the same meaning as now" you seem to mean that "the highest faculty of the mind, capacity for comprehending general truths;" is how we think of intelligence now, but the meanings of "mind", "truth" "comprehending" and "faculties of mind" have all had their own radical shifts over the last 600 years. That quoted phrase conveys an entirely different perspective and set of assumptions/implications in the context of its time, and is not at all analogous to how we read it today.

Raymond Williams' "Keywords" collects a very interesting and accessible collection of examples of this phenomenon, although it focuses more on the language of politics and society more than the language of psychology.

The modern use of intelligence, and the conceptual constellation it represents, is essentially isolated from what's described in that article, but it's re-introduction in modern psychology does borrow from its prior existence in the lexicon.

Re: Large language models lack deep insights or a theory of mind

#82

No LLMs don't think like people, they're architecturally incapable of doing so. They have, physically unlike humans no access to their own internal state and they're, save for a small context window, static systems. They also have no insights. There's a hilarious video about LLM Jailbreaks by Karpathy[1] from a week ago, where he shows how you can break model responses by asking the same question with a base64 string…

While you right about LLMs, you're not really making the case for humans well at all.

Human insight is really easy to break, confidence men wouldn't really be a thing if it were hard to break. Simply putting a statement like "I love you" in front of a statement commonly overrides our intellect. Or offering a chocolate bar in trade of our passwords. If you want a human to tell you how to end the world, you'd just convince them to be your friend first.

Re: Large language models lack deep insights or a theory of mind

#83

For me, the entire AGI conversation is hyperbolic / hype. How can we infer intelligence to something when we, ourselves, have such a poor (none) grasp of what makes us conscience? I'm associating intelligence with consciousness - because it seems correlated. Are we really ready to associate "AGI" with solving math problems ("new Q algo.")? That seems incredibly naive & reinforces my opinion that LLM's are much more l…

Completely agree, and while we are at it... look I'm just a guy, not an expert, but I can't understand why there's so much focus on AGI. It feels like there are so many niche areas where we could apply some kind of analytical augmentation and by solving problems in the small, might learn something that would help figure the larger question of intelligence. I don't need the AI to replace everything I do, I need it to…

Many of the seemingly small problems do require a good model of the world for context and edge case solving, so they still get very close to general intelegence.

Re: Large language models lack deep insights or a theory of mind

#84
post #29

Earlier quoted context omitted.

Maybe the soul is social, and oriented towards others? I believe it can be constructed. If you assume that "the eyes are the window to the soul", you notice some interesting properties. 1. It is far more observable from the outside (eyes open/lidded/closed, emotion read in eyes) 2. It affects behavior in a diffuse way 3. It pays attention but does not dictate

> Maybe the soul is social My pet theory about human consciousness is that is that consciousness is simply recursive theory of mind. Theory of mind [1] is our ability to simulate and reason about the mental states of others. It's how we predict what people are thinking and how they will react to our actions, which is critical for choosing how to act in a social environment. But when you're thinking about what's in so…

A Possible Evolutionary Reason for Why We Seem to Have Continuity of Consciousness and Personality.

Thesis : The very thing (a brain module) which allows for outside object continuity, that same brain module maintains inside self / personality / identity continuity.

Reasoning :

Evolution found out modelling the outside world is helpful for survival. Some eons later, it figured out modelling yourself (self / agent) modelling the outside world is also helpful. In the outside world, we keep track of continuity of objects through a brain module which hones in on the essence of objects (e.g. tracking a prey or predator.) so that EXACT matching algorithms aren't applied but ONLY approximate ones. As soon as a high enough approximate match (>95%) is found, we "register" it to be an exact match. i.e. the Brain bumps up the confidence level to 100%. This is also the reason why we consider our friend Bob to be the same childhood Bob even though he has different hairstyle, clothes, and other such properties. We don't call Bob who looks different than yesterday as an Imposter. The damage to this brain module could lead us to call Bob today an imposter. This same module also tracks continuity of self in a similar manner. Even though our "self" changes from childhood to adult we "register" "changing self" to be the same thing. i.e. internally we bump up the confidence to 100% when memories, etc. match and provide a coherent picture of the self. A multiple personality disorder is just different stable states of neural attractor states. Continuity is local to a personality but not global and hence transient. This local continuity of information could link up globally, giving rise to coherent single personality.

Capgras Syndrome :

Ramachandran Capgras Delusion Case : https://www.youtube.com/watch?v=3xczrDAGfT4

Re: Large language models lack deep insights or a theory of mind

#85

> A chief goal of artificial intelligence is to build machines that think like people. I disagree with the topic sentence. The goal should not be to "build machines that think like people", but to build machines that think, period. The way humans think is unlikely to be the optimal way to go about thinking anyways. Instead of talking about thinking, we should be talking about function. Less philosophy and more realit…

The problem here is this breaks the much more complicated issue of alignment.

The paperclip optimizer is a great parable here. If you build your intelligence to build as many paperclips as cheaply as possible don't be surprised when said intelligence disassembles you and the rest of the universe to do so.

So yea, HOW starts mattering a whole lot when you want to ensure it understands that it shouldn't do some particular things.

Re: Large language models lack deep insights or a theory of mind

#86
Here's my theory:

Consider a typical LLM token vector used to train and interact with an LLM.

Now imagine that other aspects of being human (sensory input, emotional input, physical body sensation, gut feelings, etc.) could be added as metadata to the the token stream, along with some kind of attention function that amplified or diminished the importance of those at any given time period -- all still represented as a stream of tokens.

If an LLM could be trained on input that was enriched by all of the above kind of data, then quite likely the output would feel much more human than the responses we get from LLMs.

Humans are moody, we get headaches, we feel drawn to or repulsed by others, we brood and ruminate at times, we find ourselves wanting to impress some people, some topics make us feel alive while others make us feel bored.

Human intelligence is always colored by the human experience of obtaining it. Obviously we don't obtain it by getting trained on terabytes of data all at once disconnected from bodily experience.

Seemingly we could simulate a "body" and provide that as real time token metadata for an LLM to incorporate, and we might get more moodiness, nostalgia, ambition, etc.

Asking for a theory of mind is in fact committing the Cartesian error of making a mind/body distinction. What is missing with LLMs is a theory of mindbody... similarity to spacetime is not accidental as humans often fail to unify concepts at first.

LLMs are simply time series predictors that can handle massive numbers of parameters in a way that allows them to generate corresponding sequences of tokens that (when mapped back into words) we judge as humanlike or intelligence-like, but those are simply patterns of logic that come from word order, which is closely related in human languages to semantics.

It's silly to think that we humans are not abstractly representable as a probabilistic time series prediction of information. What isn't?

Re: Large language models lack deep insights or a theory of mind

#87
post #78

Few weeks ago I did an experiment after a discussion here about LLMs and chess. Basically inventing a board game and play against ChatGPT and see what happened. It was not able to do a single move, even having provided all the possible start moves in the prompt as part of the rules. Not that I had a lot of hope about it, but it was definitely way worst than I expected. If someone wants to take a look at it: https://j…

I have played some moves with GPT-4 and they seem right to me, what does this mean? That the model switched from not understanding to understanding, from unintelligent to intelligent? I don't think so, GPT-4 is just a more intelligent model than GPT-3.5 and it does understand more. Also in this game if I don't move the queen I force a draw, right?

> what does this mean?

I don't know, take your own conclusions, I tried what I tried with the results I got. And the reason I created a Monte Carlo Engine to play the game was specifically because of this, I expected ChatGPT to be able to make moves but actually not being good with the game. You can try yourself, the code is available.

> Also in this game if I don't move the queen I force a draw, right?

I don't know as there is no time but I assume it is mandatory to move when, what happens in a chess game with no time if someone does not want to move? Same applies here.

Re: Large language models lack deep insights or a theory of mind

#88

For me, the entire AGI conversation is hyperbolic / hype. How can we infer intelligence to something when we, ourselves, have such a poor (none) grasp of what makes us conscience? I'm associating intelligence with consciousness - because it seems correlated. Are we really ready to associate "AGI" with solving math problems ("new Q algo.")? That seems incredibly naive & reinforces my opinion that LLM's are much more l…

[deleted]

Re: Large language models lack deep insights or a theory of mind

#89

For me, the entire AGI conversation is hyperbolic / hype. How can we infer intelligence to something when we, ourselves, have such a poor (none) grasp of what makes us conscience? I'm associating intelligence with consciousness - because it seems correlated. Are we really ready to associate "AGI" with solving math problems ("new Q algo.")? That seems incredibly naive & reinforces my opinion that LLM's are much more l…

A(G)I models don't need higher order thinking or somesuch to be impactful. For that they just need to increase productivity with or without job loss (be Good Enough), which they are on a good track for.

The real impacts will come when they are properly integrated into the current computational fabric, which everyone is racing to do as we write this.

Re: Large language models lack deep insights or a theory of mind

#90
post #55

Earlier quoted context omitted.

I think they may be referring to the principle task that consciousness serves in humans, which is to rationalize decisions we've already made subconsciously to other people so they will help us. The conscious "why" comes after the decision. In that sense it's exactly the kind of bullshit machine that LLMs are.

A thought experiment: what kind of functional MRI result would convince you that human consciousness is real and an important part of decision making? Note: if the result is someone reporting having made a decision before brain activity is seen, my next question is going to be "How does that work?"

>Note: if the result is someone reporting having made a decision before brain activity is seen, my next question is going to be "How does that work?"

This is not how we figure conscious explanations are often(always?) post-hoc rationalizations.

See -

Split brain experiments - https://www.nature.com/articles/483260a

Experiments on Choice preferences -https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3196841/

Post reply on HN