Live data from Hacker News

Is chain-of-thought AI reasoning a mirage?

seangoedecke.com

101–110 of 191 posts

Re: Is chain-of-thought AI reasoning a mirage?

#101
> but we know that reasoning is an emergent capability!

Do we though? There is widespread discussion and growing momentum of belief in this, but I have yet to see conclusive evidence of this. That is, in part, why the subject paper exists...it seeks to explore this question.

I think the author's bias is bleeding fairly heavily into his analysis and conclusions:

> Whether AI reasoning is “real” reasoning or just a mirage can be an interesting question, but it is primarily a philosophical question. It depends on having a clear definition of what “real” reasoning is, exactly.

I think it's pretty obvious that the researchers are exploring whether or not LLMs exhibit evidence of _Deductive_ Reasoning [1]. The entire experiment design reflects this. Claiming that they haven't defined reasoning and therefore cannot conclude or hope to construct a viable experiment is...confusing.

The question of whether or not an LLM can take a set of base facts and compose them to solve a novel/previously unseen problem is interesting and what most people discussing emergent reasoning capabilities of "AI" are tacitly referring to (IMO). Much like you can be taught algebraic principles and use them to solve for "x" in equations you have never seen before, can an LLM do the same?

To which I find this experiment interesting enough. It presents a series of facts and then presents the LLM with tasks to see if it can use those facts in novel ways not included in the training data (something a human might reasonably deduce). To which their results and summary conclusions are relevant, interesting, and logically sound:

> CoT is not a mechanism for genuine logical inference but rather a sophisticated form of structured pattern matching, fundamentally bounded by the data distribution seen during training. When pushed even slightly beyond this distribution, its performance degrades significantly, exposing the superficial nature of the “reasoning” it produces.

> The ability of LLMs to produce “fluent nonsense”—plausible but logically flawed reasoning chains—can be more deceptive and damaging than an outright incorrect answer, as it projects a false aura of dependability.

That isn't to say LLMs aren't useful, just exploring it's boundaries. To use legal services as an example, using an LLM to summarize or search for relevant laws, cases, or legal precedent is something it would excel at. But don't ask an LLM to formulate a logical rebuttal to an opposing council's argument using those references.

Larger models and larger training corpuses will expand that domain and make it more difficult for individuals to discern this limit; but just because you can no longer see a limit doesn't mean there is none.

And to be clear, this doesn't diminish the value of LLMs. Even without true logical reasoning LLMs are quite powerful and useful tools.

[1] https://en.wikipedia.org/wiki/Logical_reasoning

Re: Is chain-of-thought AI reasoning a mirage?

#102

> The first is that reasoning probably requires language use. Even if you don’t think AI models can “really” reason - more on that later - even simulated reasoning has to be reasoning in human language. That is an unreasonable assumption. In case of LLMs it seems wasteful to transform a point from latent space into a random token and lose information. In fact, I think in near future it will be the norm for MLLMs to "…

> In fact, I think in near future it will be the norm for MLLMs to "think" and "reason" without outputting a single "word".

It will be outputting something, as this is the only way it can get more compute - output a token, then all context + the next token is fed through the LLM again. It might not be presented to the user, but that's a different story.

Re: Is chain-of-thought AI reasoning a mirage?

#103
>but we know that reasoning is an emergent capability!

This is like saying in the 70s that we know only the US is capable of sending a man to the moon. Just because the reasoning developed in a particular context means very little about what the bare minimum requirements for that reasoning are.

Overall I am not a fan of this blogpost. It's telling how long the author gets hung up on a paper making "broad philosophical claims about reasoning", based on what reads to me as fairly typical scientific writing style. It's also telling how highly cherry-picked the quotes they criticize from the paper are. Here is some fuller context:

>An expanding body of analyses reveals that LLMs tend to rely on surface-level semantics and cluesrather than logical procedures (Chen et al., 2025b; Kambhampati, 2024; Lanham et al., 2023; Stechly et al., 2024). LLMs construct superficial chains of logic based on learned token associations, often failing on tasks that deviate from commonsense heuristics or familiar templates (Tang et al., 2023). In the reasoning process, performance degrades sharply when irrelevant clauses are introduced, which indicates that models cannot grasp the underlying logic (Mirzadeh et al., 2024)

>Minor and semantically irrelevant perturbations such as distractor phrases or altered symbolic forms can cause significant performance drops in state-of-the-art models (Mirzadeh et al., 2024; Tang et al., 2023). Models often incorporate such irrelevant details into their reasoning, revealing a lack of sensitivity to salient information. Other studies show that models prioritize the surface form of reasoning over logical soundness; in some cases, longer but flawed reasoning paths yield better final answers than shorter, correct ones (Bentham et al., 2024). Similarly, performance does not scale with problem complexity as expected—models may overthink easy problems and give up on harder ones (Shojaee et al., 2025). Another critical concern is the faithfulness of the reasoning process. Intervention-based studies reveal that final answers often remain unchanged even when intermediate steps are falsified or omitted (Lanham et al., 2023), a phenomenon dubbed the illusion of transparency (Bentham et al., 2024; Chen et al., 2025b).

You don't need to be a philosopher to realize that these problems seem quite distinct from the problems with human reasoning. For example, "final answers remain unchanged even when intermediate steps are falsified or omitted"... can humans do this?

Re: Is chain-of-thought AI reasoning a mirage?

#104
post #91

Earlier quoted context omitted.

Not all reasoning requires language. Symbolic reasoning uses language. Real-time spatial reasoning like driving a car and not hitting things does not seem linguistic. Figuring out how to rotate a cabinet so that it will clear through a stairwell also doesn't seem like it requires language, only to communicate the solution to someone else (where language can turn into a hindrance, compared to a diagram or model).

Pivot!

Can we be Friends?

Re: Is chain-of-thought AI reasoning a mirage?

#105
post #86
post #82

Earlier quoted context omitted.

Solutions to some of the hardest problems I've had have only come after a night of sleep or when I'm out on a walk and I'm not even thinking about the problem. Maybe what my brain was doing was something different from reasoning?

This is a very important point and mostly absent from the conversation. We have many words that almost mean the same thing or can mean ment different things - and conversations about intelligence and consciousness are riddled with them.

> This is a very important point and mostly absent from the conversation.

That's because when humans are mentioned at all in the context of coding with “AI”, it's mostly as bad and buggy simulations of those perfect machines.

Re: Is chain-of-thought AI reasoning a mirage?

#106

> The first is that reasoning probably requires language use. Even if you don’t think AI models can “really” reason - more on that later - even simulated reasoning has to be reasoning in human language. That is an unreasonable assumption. In case of LLMs it seems wasteful to transform a point from latent space into a random token and lose information. In fact, I think in near future it will be the norm for MLLMs to "…

> It is not a "philosophical" (by which the author probably meant "practically inconsequential") question.

I didn't take it that way. I suppose it depends on whether or not you believe philosophy is legitimate

Re: Is chain-of-thought AI reasoning a mirage?

#107

Earlier quoted context omitted.

I don't agree with the parallel. Submarines can move through water - whether you call that swimming or not isn't an interesting question, and doesn't illuminate the function of a submarine. With thinking or reasoning, there's not really a precise definition of what it is, but we nevertheless know that currently LLMs and machines more generally can't reproduce many of the human behaviours that we refer to as thinking.…

It is strange that you started your comment with "I don't agree". The rest of the comment demonstrates that you do agree.

To be more clear about why I disagree the cases are parallel:

We know how a submarine moves through water, whether it's "swimming" isn't an interesting question.

We don't know to what extent a machine can reproduce the cognitive functions of a human. There are substantive and significant questions about whether or to what extent a particular machine or program can reproduce human cognitive functions.

So I might have phrased my original comment badly. It doesn't matter if we use the word "thinking" or not, but it does matter if a machine can reproduce the human cognitive functions, and if that's what we mean by the question whether a machine can think, then it does matter.

Re: Is chain-of-thought AI reasoning a mirage?

#108

Earlier quoted context omitted.

It is strange that you started your comment with "I don't agree". The rest of the comment demonstrates that you do agree.

To be more clear about why I disagree the cases are parallel: We know how a submarine moves through water, whether it's "swimming" isn't an interesting question. We don't know to what extent a machine can reproduce the cognitive functions of a human. There are substantive and significant questions about whether or to what extent a particular machine or program can reproduce human cognitive functions. So I might have…

"We know how it moves" is not the reason the question of whether a submarine swims is not interesting. It's because the question is mainly about the definition of the word "swim" rather than about capabilities.

> if that's what we mean by the question whether a machine can think

That's the issue. The question of whether a machine can think (or reason) is a question of word definitions, not capabilities. The capabilities questions are the ones that matter.

Re: Is chain-of-thought AI reasoning a mirage?

#109

> The first is that reasoning probably requires language use. Even if you don’t think AI models can “really” reason - more on that later - even simulated reasoning has to be reasoning in human language. That is an unreasonable assumption. In case of LLMs it seems wasteful to transform a point from latent space into a random token and lose information. In fact, I think in near future it will be the norm for MLLMs to "…

You're looking at this from the perspective of what would make sense for the model to produce. Unfortunately, what really dictates the design of the models is what we can train the models with (efficiently, at scale). The output is then roughly just the reverse of the training. We don't even want AI to be an "autocomplete", but we've got tons of text, and a relatively efficient method of training on all prefixes of a sentence at the same time.

There have been experiments with preserving embedding vectors of the tokens exactly without loss caused by round-tripping through text, but the results were "meh", presumably because it wasn't the input format the model was trained on.

It's conceivable that models trained on some vector "neuralese" that is completely separate from text would work better, but it's a catch 22 for training: the internal representations don't exist in a useful sense until the model is trained, so we don't have anything to feed into the models to make them use them. The internal representations also don't stay stable when the model is trained further.

Re: Is chain-of-thought AI reasoning a mirage?

#110
post #93

Earlier quoted context omitted.

It’s incredible to me that so many seem to have fallen for “humans are just LLMs bruh” argument but I think I’m beginning to understand the root of the issue. People who only “deeply” study technology only have that frame of reference to view the world so they make the mistake of assuming everything must work that way, including humans. If they had a wider frame of reference that included, for example, Early Childhoo…

That is an issue prevalent in the western world for the last 200 years, beginning possibly with the Industrial Revolution, probably earlier. That problem is reductionism, consequently applied down to the last level: discover the smallest element of every field of science, develop an understanding of all the parts from the smallest part upwards and develop, from the understanding of the parts, an understanding of the…

Taking things apart to see how they tick is called reduction, but (re)assembling the parts is emergence.

When you reduce something to its components, you lose information on how the components work together. Emergence 'finds' that information back.

Compare differentiation and integration, which lose and gain terms respectively.

In some cases, I can imagine differentiating and integrating certain functions actually would even be a direct demonstration of reduction and emergence.

Post reply on HN