Live data from Hacker News

Large language models lack deep insights or a theory of mind

arxiv.org

131–140 of 270 posts

Re: Large language models lack deep insights or a theory of mind

#131
post #77

Earlier quoted context omitted.

I was right there with you until you mentioned Searle. :-) The Chinese room argument is bad in that it hides an assumption of mind/body dualism. If you believe that humans have "souls" and other things do not, then you have a qualitative difference between a human or a machine. On the other hand, if you are a materialist then you are faced with the problem that humans don't have much understanding of semantics either…

Searle's one of the most die-hard materialist philosophers of mind around. It's his materialism that leads him to make his argument. Computers and human brains are both made out of atoms but that doesn't mean they're not qualitatively different. By that logic I"d be no different from a tree. There's qualitative differences between computers and human brains. Our cognition is biochemical and horribly slow , just by vi…

LLMs nevertheless contain rudimentary theories, don't they? Like the Othello example demonstrating a spatial model that is emergent.

LLMs are fast like calculators but it seems the optimization process that generates the parameter weights still produces "fuzzy semantics" and the Othello emergence is just one example.

Re: Large language models lack deep insights or a theory of mind

#132

Earlier quoted context omitted.

It would have an emotional reaction to certain "thought constructs" and would be guided by that. Or we could just give them three laws

With current LLMs the three laws might be tough to implement in a way that can't be prompt injected around. That's why I described the extra bits as a "conscience" which could enforce the three laws. Maybe the three laws are the internal conscience's context prompt while the main LLM is more able to think anything in general and then the output is tuned down by the conscience? Otherwise the laws will have to be imple…

I mean, the whole purpose of the I, Robot story was to show you that the 3 laws didn't work. We had the first story on prompt injection decades ago and we just didn't realize it.

Re: Large language models lack deep insights or a theory of mind

#133
post #114

Earlier quoted context omitted.

The equivalent for a human would be an reflexive response to a question, the kind you could immediately answer after being woken up at 3am in the morning. That type of answer has been deeply trained into the human networks and also requires no deep insight. But if a human is allowed time and internal reasoning iterations, so should the LLM when determining if it has deep insight. Right now we're simply observing inpu…

Completely agree. From a computer science point of view: a single prompt/response cycle from a LLM is equivalent to a pure function; the answer is a function of the prompt and the model weights and is fundamentally reducible to solving a big math equation (in which each model parameter is a term.) It seems almost self evident that "reasoning" worthy of the name would involve some sort of iterative/recursive search pr…

Ideally a recursive execution would also be a pure function - maybe a better way to put it about current LLMs is that they are a single mathematical expression being built up from a fix number of nodes and only addition and multiplication.

Re: Large language models lack deep insights or a theory of mind

#134

Earlier quoted context omitted.

I'm older. I've bought 'new' board games for kids. Then, I have been un-able to play because the instructions were pretty bad. Humans also need to 'learn'. Need a few play-throughs. No human is going out and 'in a vacuum' with no experience, buying Risk and from scratch, read instructions and play perfect game winning strategy.

The thing is that I wanted to prove that ChatGPT was not able to learn from the rules and that is indeed a Language Model that puts one token after the other, if it know how to play chess it is because it has seen games in the past, as I say in my post: > If it is not memorizing, how do you think is doing it? (me) > by trying to learning the general rules that to explain the dataset and minimize its loss. That’s what…

> Language Model that puts one token after the other,

The interesting thing about the one-token-at-a-time process in OpenAI transformer LLMs is how the Attention Mechanism is executing ~1600 processes in parallel over the entire context window for each new token generated. So it is dynamically re-evaluating the entire context (including the rules of the game) in relation to the next token at each step.

Re: Large language models lack deep insights or a theory of mind

#136

This is a terrible eval. Do not update your beliefs on whether LLMs have Theory of Mind based on this paper. The eval is a weird, noisy visual task (picture of astronaut with “care packages”). Their results are hopelessly narrow. A better eval is to use actual scientifically tested psychology test on text (the native and strongest domain for LLMs), for example the sort of scenarios used to gauge when children develop…

Or there are enough of those examples in the training set that it can guess well. Not sure how such an example would prove anything when we know an LLM is just guessing the best words. Nothing I’ve seen shows evidence of any sort of abstract concepts in there.

Wouldn't this also be the same for humans?

Re: Large language models lack deep insights or a theory of mind

#137
post #49

Earlier quoted context omitted.

Most of what you're saying here is describing the alignment issue. We (mostly) don't want unaligned A(G|S)I. The outcomes of that could be extenstential.

Only for those mundane senses of alignment where we say "This system is reliable in tasks that look like X and unreliable in tasks that look like Y, so let's craft hard boundaries to avoid naive use for Y" But it's skeptical of the other sense alignment, where a potential Master Strategist needs to be trained or crippled before it outsmarts us. It sees that perspective as comparable to logicians debating whether we m…

" but history and analysis give room for skeptics to be like

When the skeptic is correct. The problem with skeptics is when incorrectness is not terminal, they can't hear you over the sound of pushing the goal posts farther to give a reasonable rebuttal for their originally incorrect statements.

Re: Large language models lack deep insights or a theory of mind

#138
post #83

Earlier quoted context omitted.

Completely agree, and while we are at it... look I'm just a guy, not an expert, but I can't understand why there's so much focus on AGI. It feels like there are so many niche areas where we could apply some kind of analytical augmentation and by solving problems in the small, might learn something that would help figure the larger question of intelligence. I don't need the AI to replace everything I do, I need it to…

Many of the seemingly small problems do require a good model of the world for context and edge case solving, so they still get very close to general intelegence.

Yep, at least in my eyes you'll never be able to "solve" self driving without solving the G in AGI. You require a world model for predictions in order to have enough time to avoid many bad outcomes. Avoiding an empty soda can and avoiding a brick are similar problems, but one can easily lead to critical failures if you miss it.

Re: Large language models lack deep insights or a theory of mind

#139
post #130

Earlier quoted context omitted.

There are plenty of models that use introspection and check answers, that’s the idea behind let’s think about it step by step.

I feel the "let's think about it step by step" is a bit of a hack. To circumvent the fact that there's no external loop you use the fact that it gets re-run on every token so you can store a bit of state in the tokens that it's already generated. Or am I misunderstanding something about that technique?

You are right, it’s sometimes called zero shot chain of thought, but it’s a way of getting the type of thing you are describing to happen. The LLMs somehow process things in a perceived step by step to get a much improved answer. Whether the external loop or an llm imposed internal loop, does it matter? Are our own minds looping or just adding tokens?

Re: Large language models lack deep insights or a theory of mind

#140

For me, the entire AGI conversation is hyperbolic / hype. How can we infer intelligence to something when we, ourselves, have such a poor (none) grasp of what makes us conscience? I'm associating intelligence with consciousness - because it seems correlated. Are we really ready to associate "AGI" with solving math problems ("new Q algo.")? That seems incredibly naive & reinforces my opinion that LLM's are much more l…

Completely agree, and while we are at it... look I'm just a guy, not an expert, but I can't understand why there's so much focus on AGI. It feels like there are so many niche areas where we could apply some kind of analytical augmentation and by solving problems in the small, might learn something that would help figure the larger question of intelligence. I don't need the AI to replace everything I do, I need it to…

>I need it to solve 10,000 micro problems I solve every day - each of which is a business opportunity for someone.

Because you have to solve 10,000 different problems. And a huge number of those problems are going to have significant overlap, but sharing lessons between them is going to be difficult unless you have a generalized algorithm.

Hence AGI is the trillion dollar question.

Post reply on HN