Earlier quoted context omitted.
Spicy autocomplete is still spicy autocomplete
I'm not sure if system capable of ie. reasoning over images deserves this label anymore?
The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]
181–190 of 276 posts
Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]
#182Earlier quoted context omitted.
> Rather, the problem is, once you do define it, you will quickly find that LLMs are capable of it. That’s really not what’s happening though. People who claim LLMs can do X and Y often don’t even understand how LLMs work. The opposite is also true. They just open a prompt and get an output and shout Eureka. Of course not everyone is like this, but majority are. It’s similar to what we think about thinking itself. Yo…
Notice how nearly every comment is just dancing around the issue and nitpicking instead of just owning up to whether or not they think LLMs are capable of reasoning.
I personally believe that not being sure about something, esp. on a topic as complicated and as this one, remaining open to different possibilities, and having a bit of skepticism, is healthy.
Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]
#183Earlier quoted context omitted.
> Humans invented machines because they could not do certain things. If your disappointment is that the LLM didn't invent a computer to solve the problem, maybe you need to give it access to physical tools, robots, labs etc.
Nah, even if we follow such a weak "argument" the fact is that, ironically, the evidence shown in this and other papers point towards the idea that even if LRMs did have access to physical tools, robots labs, etc*, they probably would not be able to harness them properly. So even if we had an API-first world (i.e. every object and subject in the world can be mediated via a MCP server), they wouldn't be able to perfor…
If your argument is just that LRMs are more noisy and error prone in their reasoning, then I don't disagree.
> I'm kind of tired of people comparing humans to machines in such simple and dishonest ways.
The issue is people who say "see, the AI makes mistakes at very complex reasoning problems, so their 'thinking is an illusion'". That's the title of the paper.
This mistake comes not from people "comparing humans to machines", but from people fundamentally misunderstanding what thinking is. If thinking is what humans do, then errors are expected.
There is this armchair philosophical idea, that a human can simulate any turning machine and thus our reasoning is "maxomally general", and anything that can't do this is not general intelligence. But this is the complete opposite of reality. In our world, anything we know that can perfectly simulate a turning machine is not general intelligence, and vice versa.
Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]
#184All the environments the test (Tower of Hanoi, Checkers Jumping, River Crossing, Block World) could easily be solved perfectly by any of the LLMs if the authors had allowed it to write code. I don't really see how this is different from "LLMs can't multiply 20 digit numbers"--which btw, most humans can't either. I tried it once (using pen and paper) and consistently made errors somewhere.
The goal isnt to assess the LLM capability at solving any of those problems. The point isnt how good they are at block world puzzles. The point is to construct non-circular ways of quantifying model performance in reasoning. That the LLM has access to prior exemplars of any given problem is exactly the issue in establishing performance in reasoning, over historical synthesis.
Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]
#185Earlier quoted context omitted.
> I don't really see how this is different from "LLMs can't multiply 20 digit numbers"--which btw, most humans can't either. I tried it once (using pen and paper) and consistently made errors somewhere. People made missiles and precise engineering like jet aircraft before we had computers, humans can do all of those things reliably just by spending more time thinking about it, inventing better strategies and using mo…
Some specialized people could probably do 20x20, but I'd still expect them to make a mistake at 100x100. The level we needed for space crafts was much less than that, and we had many levels of checks to help catch errors afterwards. I'd wager that 95% of humans wouldn't be able to do 10x10 multiplication without errors, even if we paid them $100 to get it right. There's a reason we had to invent lots of machines to h…
With enough effort and time we can arrive at a perfect solution to those problems without a computer.
This is not a hypothetical, it was like that for at least hundreds of years.
Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]
#186In figure 1 bottom-right they show how the correct answers are being found later as the complexity goes higher. In the description they even state that in false responses the LRM often focusses on a wrong answer early and then runs out of tokens before being able to self-correct. This seems obvious and indicates that it’s simply a matter of scaling (bigger token budget would lead better abilities for complexer tasks)…
Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]
#187Earlier quoted context omitted.
Notice how nearly every comment is just dancing around the issue and nitpicking instead of just owning up to whether or not they think LLMs are capable of reasoning.
I wish that was the case. I see most people, who believe to be experts, already made up their minds and being pretty sure about it. I personally believe that not being sure about something, esp. on a topic as complicated and as this one, remaining open to different possibilities, and having a bit of skepticism, is healthy.
Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]
#188Earlier quoted context omitted.
Reasoning exists on a spectrum, not as a binary property. I'm not claiming that LLMs reason identically to humans in all contexts. You act as if statistical processes can’t ever scale into reasoning, despite the fact that humans themselves are gradient-trained statistical learners over evolutionary and developmental timescales.
I think you might re-read what I wrote. I very strongly agree with you. :)
Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]
#189Earlier quoted context omitted.
Nah, even if we follow such a weak "argument" the fact is that, ironically, the evidence shown in this and other papers point towards the idea that even if LRMs did have access to physical tools, robots labs, etc*, they probably would not be able to harness them properly. So even if we had an API-first world (i.e. every object and subject in the world can be mediated via a MCP server), they wouldn't be able to perfor…
> Don't misinterpret me, human errors do happen in those contexts because, well, we're talking about humans, but not as catastrophically as the errors committed by LRMs in this paper. If your argument is just that LRMs are more noisy and error prone in their reasoning, then I don't disagree. > I'm kind of tired of people comparing humans to machines in such simple and dishonest ways. The issue is people who say "see,…
That's not what the paper proposes (i.e. it commits errors => thinking is an illusion). It in fact looks at the failures modes and then it argues that due to HOW they fail and in which contexts/conditions, that their thinking may be "illusory" (not that the word illusory matters that much, papers of this calibre always strive for interesting sounding titles). Hell, they even gave the exact algo to the LRM, it probably can't get more enabling than that.
Humans are lossy thinkers and error-prone biological "machines", but an educated+aligned+incentivized one shouldn't have problems following complex instructions/algos (not in a no-errors way, but rather, in a self-correcting way); we thought that LRMs did that too, but the paper shows how they even start using less "thinking" tokens after a complexity threshold and that's terribly worrisome, akin to someone getting frustrated and stopping thinking after a problem gets too difficult which goes contrary to the idea that these machines can run laboratories by themselves. It is not the last nail in the coffin because more evidence is needed as always, but when taken into account with other papers, it points towards the limitations of LLMs/LRMs and how those limitations may not be solvable with more compute/tokens, but rather exploring new paradigms (long due in my opinion, the industry usually forces a paradigm as panacea during hype cicles in the name of hypergrowth/sales).
In short the argument you say the paper and posters ITT make is very different from what they are actually saying, so beware of the logical leap you are making.
> There is this armchair philosophical idea, that a human can simulate any turning machine and thus our reasoning is "maxomally general", and anything that can't do this is not general intelligence. But this is the complete opposite of reality. In our world, anything we know that can perfectly simulate a turning machine is not general intelligence, and vice versa.
That's typical goalpoast moving and happens in both ways when talking about "general intelligence" as you say, since the dawn of AI and the first neural networks. I'm not following why this is relevant for the discussion though.
Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]
#190Earlier quoted context omitted.
The goal isnt to assess the LLM capability at solving any of those problems. The point isnt how good they are at block world puzzles. The point is to construct non-circular ways of quantifying model performance in reasoning. That the LLM has access to prior exemplars of any given problem is exactly the issue in establishing performance in reasoning, over historical synthesis.
How are these problems more interesting than simple arithmetic or algorithmic problems?