Live data from Hacker News

Large language models lack deep insights or a theory of mind

arxiv.org

91–100 of 270 posts

Re: Large language models lack deep insights or a theory of mind

#91

For me, the entire AGI conversation is hyperbolic / hype. How can we infer intelligence to something when we, ourselves, have such a poor (none) grasp of what makes us conscience? I'm associating intelligence with consciousness - because it seems correlated. Are we really ready to associate "AGI" with solving math problems ("new Q algo.")? That seems incredibly naive & reinforces my opinion that LLM's are much more l…

Completely agree, and while we are at it... look I'm just a guy, not an expert, but I can't understand why there's so much focus on AGI. It feels like there are so many niche areas where we could apply some kind of analytical augmentation and by solving problems in the small, might learn something that would help figure the larger question of intelligence. I don't need the AI to replace everything I do, I need it to…

I've thought about this a bit as well, and I think it's almost like this toxic concoction of incentive (how can "we" hype this until and make boatloads of money off of it, coupled with a genuine (if sub-conscious) desire to be seen as a visionary/great engineer who "created artificial life." I mean, at least on HN, I see lots of this aspirational attitude for living the sci-fi future circa. Star trek, ex-machina, etc, while couching their language in professions of expertise now that the firehose of cash has turned on.

Also there is the general hubris in all this to only look at the new and shiny, I remember when there was that pizza robot (some multi-dimension axis hand thing) that cost whatever in building and research, when the costco pizza "robot" is pretty darn good, but doesn't sell as "futuristic/cool" because its a spigot on a servo.

Re: Large language models lack deep insights or a theory of mind

#92

For me, the entire AGI conversation is hyperbolic / hype. How can we infer intelligence to something when we, ourselves, have such a poor (none) grasp of what makes us conscience? I'm associating intelligence with consciousness - because it seems correlated. Are we really ready to associate "AGI" with solving math problems ("new Q algo.")? That seems incredibly naive & reinforces my opinion that LLM's are much more l…

A particular subset of Connectivism have a philosophical belief that the mind IS a neutral net, not that it is a reductive practical model.

Hinton is one of these individuals and with no definition of what intelligence is it is an understandable of dogmatic position.

This whole problem of not being able to define what intelligence is pretty much allows us all to pick and choose.

In my mind BPP is the complexity class solvable by ANNs and it is a safe and educated guess that most likely BPP=P.

BPP being one of the largest practical complexity classes makes work in this area valuable.

But due to many reasons that I won't enumerate again AGI simply isn't possible and requires a dogmatic position to believe in for people who have even a basic understanding of how they work and the limits from the work of Gödel etc...

But many of the top scientists in history have been believers of numerology etc...

Associating math with LLMs is a useful too to avoid wasted effort by those who don't believe AGI is close, but it won't convince those who are true believers.

LLM's are very useful for searching very large dimensional spaces and for those problems that are ergotic with the Markov property they can find real answers.

But for most of what is popular in the press will almost certainly be a dead end for generalized use of the systems are not extremely error tolerant.

Unfortunately it may take another AI winter to break the hype train but I hope not.

IMHO it will have a huge impact but overconfident claims will cause real pain and misapplication for the foreseeable future.

Re: Large language models lack deep insights or a theory of mind

#93

For me, the entire AGI conversation is hyperbolic / hype. How can we infer intelligence to something when we, ourselves, have such a poor (none) grasp of what makes us conscience? I'm associating intelligence with consciousness - because it seems correlated. Are we really ready to associate "AGI" with solving math problems ("new Q algo.")? That seems incredibly naive & reinforces my opinion that LLM's are much more l…

Couldn't agree more. How about this -- I think we've already reached AGI. Let me know if this tracks: Pick a set of tasks that can be considered AGI tasks. Provided the task sequences can be compared as closer to AGI or further from AGI, we can create a reward model using the same techniques as were used by ChatGPT via RLHF. Thus, for any definition of AGI that is meaningful and selectable, even if subjectively selectable or arbitrarily preferential, we can create a reward model for it.

You might say, well thats not AGI, AGI must also do such and such. Well, we can get arbitrarily close to that definition as well via RLHF.

Another objection might be: well, if thats the definition of AGI, that seems really underwhelming compared to the hype train. This says nothing about autonomy, sentience, free will -- exactly. Those concepts can or should be orthogonal to doing productive work.l, IMHO.

So, there it is. We can now make a reward model for folding socks, and use gradient descent with RL to do the motion planning.

Maybe thats AGI and maybe its not, but I'd really love it if we had a golden period between now and total enshittification that involved laundry folding robots.

Re: Large language models lack deep insights or a theory of mind

#94
post #42

I have small kids, toddlers, who can already speak the language but still developing their "sense of the world" or "theory of mind" if you will. Maybe it's just me, but talking to toddlers often reminds me of interacting with LLMs, where you would have this realization from time to time "oh, they don't get this, need to break down more to explain". Of course LLM has more elaborate language skills due to its exposure…

It's not just true about toddlers but also for adults in particular time frame. Maturity of thought is cultural phenomenon. Descartes used to think animals are automaton while they behaved exactly like humans in almost all aspects in which he could investigate animals and humans during those times and yet he reached illogical conclusion.

That's a great point. Just thinking out loud, if we can time travel back to the cavemen time, and assuming we speak their language, there would still be so much that we couldn't explain or they wont' be able to understand even for the smartest cavemen adults. Unless, of course we spend significant time and effort to "bring them up to speed" with modern education.

Re: Large language models lack deep insights or a theory of mind

#95

Earlier quoted context omitted.

You haven't specified what model did you use, and the green ChatGPT icon in the shared conversation usually signifies GPT-3.5 model. Here's my attempt at similar conversation — it seems GPT-4 is able to visualise the board and at least do a valid first move. https://chat.openai.com/share/98427e21-678c-4290-aa8f-da8e93...

Interesting. The model was whatever was up that that time, so probably was 3.5 if you say so.

Your conversation is from June, GPT-4 was available for almost half a year at that point.

Re: Large language models lack deep insights or a theory of mind

#96
post #78

Earlier quoted context omitted.

I have played some moves with GPT-4 and they seem right to me, what does this mean? That the model switched from not understanding to understanding, from unintelligent to intelligent? I don't think so, GPT-4 is just a more intelligent model than GPT-3.5 and it does understand more. Also in this game if I don't move the queen I force a draw, right?

> what does this mean? I don't know, take your own conclusions, I tried what I tried with the results I got. And the reason I created a Monte Carlo Engine to play the game was specifically because of this, I expected ChatGPT to be able to make moves but actually not being good with the game. You can try yourself, the code is available. > Also in this game if I don't move the queen I force a draw, right? I don't know…

>I don't know, take your own conclusions

The API cost of the game is getting noticeable, but I think you were just being naive about LLM limitations, there is simply no way that it can answer all questions simply by memorization. A simpler way is to just invent a programming language and ask the model to solve problems with it, at least I don't have to write down the position of a game ahah

Also I have trained models to do additions in the past, I removed many possible combinations of digits to show in the dataset, but after training the model was able to solve all of them, meaning it learned the algorithm and not just memorized the answers, I did it because a friend of mine thought like you that LLMs just memorize the answers from the dataset and cannot learn, but that is not how they work.

About the game, I realize that I cannot move the queen like in chess, so in this game I will eventually fall into a zugzwang, trying not to move the queen.

Re: Large language models lack deep insights or a theory of mind

#97

Earlier quoted context omitted.

There does seem to be a general factor of intelligence in humans though that is the single biggest indicator of performance. Yes there are other factors too. >Here, the authors point out that the current batch of programs are not good at tasks that benefit from a theory of mind. Not good at tasks that benefit from a theory of mind extracted from visual data.

"Seem" is doing a lot of work here. So is the implicit claim that theory of mind in general can be demonstrated by current-gen foundation models, and only those aspects dependent on vision cannot.

"Seem" only seems to be doing lots of work. Here's a relevant article: https://en.wikipedia.org/wiki/G_factor_(psychometrics)

Re: Large language models lack deep insights or a theory of mind

#98

Another paper in a long series that confuses "our tests against currently available LLMs tuned for specific tasks found that they didn't perform well on our task" with "LLMs are architecturally unsuitable for our task".

Our tests against current cars found that they didn't perform well on transatlantic flights... But who knows what the future holds? Maybe we should test them again next year. LLM names an specific product, aimed at solving an specific problem.

our tests of combustion engine driven crankshafts connected to a spinning mechanism can perform well in both transatlantic flights and cross country road trips.

Re: Large language models lack deep insights or a theory of mind

#99
post #30

I think that if they would, that would be very surprising and indicative of a lot of wastefulness inside the model architecture. All these tests are simple single prompt experiments, so the LLM's get no chance to reason about their responses. They're just system 1 thinking, the equivalent of putting a gun to someone's head and asking them to solve a large division in 2 seconds. I bet a lot of these experiments would…

Just note that loop doesn't have to be visible from outside. It can be internal, with another driving thread asking right questions. Inner monologue. Then the summary is given back to user. This will give the model space for 'thinking' with internally generated text much large than the visible prompt + output. This way multi-step logic can be implemented.

Re: Large language models lack deep insights or a theory of mind

#100
post #30

I think that if they would, that would be very surprising and indicative of a lot of wastefulness inside the model architecture. All these tests are simple single prompt experiments, so the LLM's get no chance to reason about their responses. They're just system 1 thinking, the equivalent of putting a gun to someone's head and asking them to solve a large division in 2 seconds. I bet a lot of these experiments would…

Yep, prototype exactly that this past week. With a strong instruction spec prompt from the start, you can have an AI come up with a much better answer by making sure it knows it has time to answer the questions and how it should approach the problem in stages.

The great part is with clear enough directions it also knows how to evaluate whether its done or not.

Post reply on HN