Live data from Hacker News

Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens

arstechnica.com

81–90 of 140 posts

Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens

#81
post #38

Earlier quoted context omitted.

I was just at the KDD conference and the general consensus agreed with this paper. There was only one keynoter who just made the assumption that LLMs are associated with reasoning, which was jarring as the previous keynoter had just explained at length why we need a neuro-symbolic approach instead. The thing is, I think the current companies making LLMs are _not_ trying to be correct or right. They are just trying to…

"as the previous keynoter had just explained at length why we need a neuro-symbolic approach instead" Do you have a link to the video for that talk ?

I don't think they were recorded. In fact, I don't think any of KDD gets recorded.

I think it was Dan Roth who talked about the challenges of reasoning from just adding more layers and it was Chris Manning who just quickly mentioned at the beginning of his talk that LLMs were well known for reasoning.

https://kdd2025.kdd.org/keynote-speakers/

Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens

#82
post #54

Earlier quoted context omitted.

I was just at the KDD conference and the general consensus agreed with this paper. There was only one keynoter who just made the assumption that LLMs are associated with reasoning, which was jarring as the previous keynoter had just explained at length why we need a neuro-symbolic approach instead. The thing is, I think the current companies making LLMs are _not_ trying to be correct or right. They are just trying to…

As to general consensus, Hinton gave a recent talk, and he seemed adamant that neural networks (which LLMs are) really are doing reasoning. He gives his reasons for it. Is Hinton considered an outlier or?

A) Hinton is quite vocal about desiring to be an outsider/outlier as he says it is what lets him innovate.

B) He is also famous for his Doomerism, which often depends on machines doing "reasoning".

So...it's complicated, and we all suffer from confirmation bias.

Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens

#83

Earlier quoted context omitted.

I have nothing against researching this, I think it's important. My main issue is with articles choosing to grab a "conclusion" and imply it extrapolates to larger models, without any support for that. They are going for the catchy title first, fine-print be damned.

I was just at the KDD conference and the general consensus agreed with this paper. There was only one keynoter who just made the assumption that LLMs are associated with reasoning, which was jarring as the previous keynoter had just explained at length why we need a neuro-symbolic approach instead. The thing is, I think the current companies making LLMs are _not_ trying to be correct or right. They are just trying to…

> None of this needs the model to grow strong reasoning skills. That's not where the real money is.

I never thought about it like that, but it sounds plausible.

However, I feel like getting to this stage is even harder to get right compared to reasoning?

Aside from the permanence and memory

They're currently essentially stateless, while that's surely enough for short term attachment, I'm not seeing this becoming a bigger issue because if that glaring shortfall.

It'd be like being in a relationship with a person with dementia, thats not a happy state of being.

Honestly, I think this trend is severely overstated until LLMs can sufficiently emulate memories and shared experiences. And that's still fundamentally impossible, just like "real" reasoning with understanding.

So I disagree after thinking about it more - emulated reasoning will likely have a bigger revenue stream via B2E applications compared to emotional attachment in B2C...

Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens

#84
post #42

“ the researchers created a carefully controlled LLM environment in an attempt to measure just how well chain-of-thought reasoning works when presented with "out of domain" logical problems that don't match the specific logical patterns found in their training data.” Why? If it’s out of domain we know it’ll fail.

Its getting to the nub of whether models can extrapolate instead of interpolate.

If they had _succeeded_, we'd all be taking it as proof that LLMs can reason, right?

Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens

#85
We're rapidly reaching trough of disillusionment with LLMs, and other generative transformer models for that matter. I am happy because it will help a lot of misinformed people understand what is and isn't possible (100+% productivity gains are not).

Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens

#86
post #85

We're rapidly reaching trough of disillusionment with LLMs, and other generative transformer models for that matter. I am happy because it will help a lot of misinformed people understand what is and isn't possible (100+% productivity gains are not).

don't know, maybe in the tecnical circles, but for users the thrill is still going on, and rising

Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens

#87
post #85

We're rapidly reaching trough of disillusionment with LLMs, and other generative transformer models for that matter. I am happy because it will help a lot of misinformed people understand what is and isn't possible (100+% productivity gains are not).

100% productivity gains on coding tasks are absolutely within the realm of possibility

Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens

#88

Earlier quoted context omitted.

> conclusions gathered from toy models and implying this generalises to production LLMs is useless You are just trotting out the tired argument that model size magically fixes the issues, rather than just improves the mirage, and so nothing can be known about models with M parameters by studying models with N Given enough parameters, a miraculous threshold is reached whereby LLMs switch from interpolating to extrapol…

That’s what has been seen in practice though. SOTA LLMs have been shown again and again to solve problems unseen in their data set; and despite their shortcomings they have become extremely useful for a wide variety of tasks.

> SOTA LLMs have been shown again and again to solve problems unseen in their data set

We have no idea what the training data is though, so you can't say that.

> and despite their shortcomings they have become extremely useful for a wide variety of tasks.

That seems like a separate question.

Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens

#89

Earlier quoted context omitted.

Not sure why everyone is downvoting you as I think you raise a good point - these anthropomorphic words like "reasoning" are useful as shorthands for describing patterns of behaviour, and are generally not meant to be direct comparisons to human cognition. But it goes both ways. You can still criticise the model on the grounds that what we call "reasoning" in the context of LLMs doesn't match the patterns we associat…

""Sam Altman says the perfect AI is “a very tiny model with superhuman reasoning"."" It is being marketed as directly related to human reasoning.

Sure, two things can be true. Personally I completely ignore anything Sam Altman (or other AI company CEOs/marketing teams for that matter) says about LLMs.

Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens

#90
post #78

Earlier quoted context omitted.

Mind linking any examples (or categories) of problems that are definitively not in pre training data but can still be solved by LLMs? Preferably something factual rather than creative, genuinely curious. Dumb question but anything like this that’s written about on the internet will ultimately end up as training fodder, no?

How about the International Math Olympiad? https://arstechnica.com/ai/2025/07/google-deepmind-earns-gol...

You're saying they don't use math textbooks and math forums to train LLMs, then?
Post reply on HN