Earlier quoted context omitted.
In this case it found generic advice and was confusing itself.
That's one explanation. Another could be that it simply has no real *understanding* of anything. It simply did a statistical comparison of the question to the available advice and picked the best match --- kinda what a search engine might do. Expecting *understanding* from a synthetic, statistical process will often end in disappointment.
The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]
221–230 of 276 posts
Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]
#222I wonders if there’s past symbolic reasoning research we could integrate into LLMs. They’re really good at parsing text and understanding the relationships between objects ie getting the “symbols” correct. Maybe we plug into something like prolog (or other such strategies?)
Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]
#223All "reasoning" models hit a complexity wall where they completely collapse to 0% accuracy. No matter how much computing power you give them, they can't solve harder problems. This research suggests we're not as close to AGI as the hype suggests. Current "reasoning" breakthroughs may be hitting fundamental walls that can't be solved by just adding more data or compute. Apple's researchers used controllable puzzle env…
I think this might be part of the reason Apple is “behind” on generative AI … LLMs have not really proven to be useful outside of relatively niche areas such as coding assistants, legal boiler plate and research, and maybe some data science/analysis which I’m less familiar with Other “end user” facing use cases have so far been comically bad or possibly harmful, and they just don’t meet the quality bar for inclusion…
None of them are doing the equivalent of “vibe-coding”, but they use LLMs to get 20-50% done, then take over from there.
Apple likes to deliver products that are polished. Right now the user needs to do the polishing with LLM output. But that doesn’t mean it isn’t useful today
Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]
#224All "reasoning" models hit a complexity wall where they completely collapse to 0% accuracy. No matter how much computing power you give them, they can't solve harder problems. This research suggests we're not as close to AGI as the hype suggests. Current "reasoning" breakthroughs may be hitting fundamental walls that can't be solved by just adding more data or compute. Apple's researchers used controllable puzzle env…
No matter how much computing power you give them, they can't solve harder problems. Why would anyone ever expect otherwise? These models are inherently handicapped and always will be in terms of real world experience. They have no real grasp of things people understand intuitively --- like time or money or truth ... or even death. They only *reality* they have to work from is a flawed statistical model built from the…
Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]
#225Interestingly I just hit an example of this. Highly specific but I was asking about pickleball strategy and grok and Claude both couldn’t seem to understand you can’t aim at the opponent’s feet when you’re hitting up. Just kept regurgitating internet advice and I couldn’t get it to understand the reasoning on why it was wrong.
Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]
#226My read of this is that the paper demonstrates that given a particular model (and the problems examined with it) that giving more thought tokens does not help on problems above a certain complexity. It does not say anything about the capabilities of future, larger, models to handle more complex tasks. (NB: humans trend similarly)
My concern is that people are extrapolating from this to conclusions about LLM's generally, and this is not warranted
The only part about this i find even surprising is he abstract's conclusion (1): that 'thinking' can lead to worse outcomes for certain simple problem. (again though, maybe you can say humans are the same here. You can overthink things)
Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]
#227Earlier quoted context omitted.
No matter how much computing power you give them, they can't solve harder problems. Why would anyone ever expect otherwise? These models are inherently handicapped and always will be in terms of real world experience. They have no real grasp of things people understand intuitively --- like time or money or truth ... or even death. They only *reality* they have to work from is a flawed statistical model built from the…
I agree, obviously, but half the internet is still running around claiming we're on the verge of a singularity, so demonstrating the actual limitations of these systems concretely is important.
Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]
#228Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]
#229Earlier quoted context omitted.
>There's nothing "omniscient" or "dim-witted" about these tools I disagree in that that seems quite a good way of describing them. All language is a bit inexact. Also I don't buy we are no closer to AI than ten years ago - there seem lots going on. Just because LLMs are limited doesn't mean we can't find or add other algorithms - I mean look at alphaevolve for example https://www.technologyreview.com/2025/05/14/11164…
> I figure it's hard to argue that that is not at least somewhat intelligent? The fact that this technology can be very useful doesn't imply that it's intelligent. My argument is about the language used to describe it, not about its abilities. The breakthroughs we've had is because there is a lot of utility from finding patterns in data which humans aren't very good at. Many of our problems can be boiled down to this…
Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]
#230Earlier quoted context omitted.
That's one explanation. Another could be that it simply has no real *understanding* of anything. It simply did a statistical comparison of the question to the available advice and picked the best match --- kinda what a search engine might do. Expecting *understanding* from a synthetic, statistical process will often end in disappointment.
Yup. It’s time for us to just accept that LLMs are “similar in meaning” machines not “thinking/ understanding” machines
And the same applies to a lot of real world situations.