Earlier quoted context omitted.
As to general consensus, Hinton gave a recent talk, and he seemed adamant that neural networks (which LLMs are) really are doing reasoning. He gives his reasons for it. Is Hinton considered an outlier or?
A) Hinton is quite vocal about desiring to be an outsider/outlier as he says it is what lets him innovate . B) He is also famous for his Doomerism, which often depends on machines doing "reasoning". So...it's complicated, and we all suffer from confirmation bias.
Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
131–140 of 140 posts
Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
#132Earlier quoted context omitted.
As to general consensus, Hinton gave a recent talk, and he seemed adamant that neural networks (which LLMs are) really are doing reasoning. He gives his reasons for it. Is Hinton considered an outlier or?
Link to the talk?
Ultimately I somehwat disagreed with some of Hintons points in this talk, and after some thought I came up with specific reasons/doubts, and yet at the same time, his intuitive explanations helped shift my views somewhat as well.
Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
#133Earlier quoted context omitted.
I think this policy was, in large part, intended to respect the user base, who get exhausted answering the same question over and over. I do agree they later trashed that relationship with the Monica incident and AI policies.
Sounds like they optimised for a select 1% class of self appointed gatekeepers rather than the broad user base. Classic mistake of nearly every defunct social site.
It worked beautifully for quite a while. I don't think anyone anticipated ChatGPT when planning it all out.
Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
#134Earlier quoted context omitted.
Because model size is a trivial parameter, and not a new paradigm. What you're saying is like, you can't extrapolate that long division works on 100 digit numbers because you only worked through it using 7 digit numbers and a few small polynomials.
Scale changes the performance of LLMs. Sometimes, we go so far as to say there is "emergence" of qualitative differences. But really, this is not necessary (and not proven to actually occur). What is true is that the performance of LLMs at OOD tasks changes with scale. So no, it's not the same as solving a math problem.
Of course performance improves on the same tasks.
The researchers behind the submitted work chose a certain size and certain size problems, controlling everything. There is no reason to believe that their results won't generalize to larger or smaller models.
Of course, not for the input problems being held constant! That is as strawman.
Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
#135Earlier quoted context omitted.
> Then when it fails to apply the "reasoning", that's evidence the artificial expertise we humans perceived or inferred is actually some kind of illusion. That doesn't follow, if the weakness of the model manifests on a different level we wouldn't call rational in a human. For example, a human might have dyslexia, a disorder on the perceptive level. A dyslexic can understand and explain his own limitation, but that d…
Typically when a human has a disorder or limitation they adapt to it by developing coping strategies or making use of tools and environmental changes to compensate. Maybe they expect a true reasoning model to be able to do the same thing?
It's a bit like asking human to read text and guess gender or emotional state of the author who wrote it. You just don't have this information.
Similarly you could ask why ":) is smiling and :D is happy" where the question will be seen as "[50372, 382, 62529, 326, 712, 35, 382, 7150]" - encoding looses this information, it's only visible in image rendering of this text.
Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
#136Earlier quoted context omitted.
Typically when a human has a disorder or limitation they adapt to it by developing coping strategies or making use of tools and environmental changes to compensate. Maybe they expect a true reasoning model to be able to do the same thing?
The argument is that letter level information is something llms don't have a chance to see. It's a bit like asking human to read text and guess gender or emotional state of the author who wrote it. You just don't have this information. Similarly you could ask why ":) is smiling and :D is happy" where the question will be seen as "[50372, 382, 62529, 326, 712, 35, 382, 7150]" - encoding looses this information, it's o…
The point is that if the model were really "reasoning", it would fail differently. Instead, what happens is consistent with it BSing on a textual level.
Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
#137Earlier quoted context omitted.
> SOTA LLMs have been shown again and again to solve problems unseen in their data set We have no idea what the training data is though, so you can't say that. > and despite their shortcomings they have become extremely useful for a wide variety of tasks. That seems like a separate question.
I have applied O3 pro on unpublished abandoned research of mine that was never published and lives in an intersection that is as entirely novel as it's uninteresting. O3 pro (but not O3) was successfully able to apply reasoning and math to this domain in interesting ways, much like an expert researcher in these areas would. Again, the field and the problem is with 100% certainty OOD of the data. However, the techniqu…
Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
#138Earlier quoted context omitted.
I was just at the KDD conference and the general consensus agreed with this paper. There was only one keynoter who just made the assumption that LLMs are associated with reasoning, which was jarring as the previous keynoter had just explained at length why we need a neuro-symbolic approach instead. The thing is, I think the current companies making LLMs are _not_ trying to be correct or right. They are just trying to…
As to general consensus, Hinton gave a recent talk, and he seemed adamant that neural networks (which LLMs are) really are doing reasoning. He gives his reasons for it. Is Hinton considered an outlier or?
I recently had fun asking Gemini to compare how Wittgenstein and Chomsky would view calling a large transformer that was trained entirely on a synthetic 'language' (in my case symbols that encode user behaviour in an app) a 'language' or not. And then, for the killer blow, whether an LLM that is trained on Perl is a language model.
My point being that whilst Hinton is a great and all, I don't think I can quite pin down his definitions of the precise words like reasoning etc. Its possible for people to have opposite meanings for the same words (Wittgenstein famously had two contradictory approaches in his lifetime). In the case of Hinton, I can't quite pin down how loosely or precisely he is using the terms.
A forward-only transformer like GPT can only do symbolic arithmetic to the depth of its layers, for example. And I don't think the solution is to add more layers.
Of course humans are entirely neuro and we somehow manage to 'reason'. So YMMV.
Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
#139Earlier quoted context omitted.
> It really makes me believe that the models do not really understand the topic, even the basics but just try to predict the text. This is correct. There is no understanding, there aren't even concepts. It's just math, it's what we've been doing with words in computers for decades, just faster and faster. They're super useful in some areas, but they're not smart, they don't think.
I’ve never seen so much misinformation trotted out by the laity as I have with LLMs. It’s like I’m in a 19th century forum with people earnestly arguing that cameras can steal your soul. These people haven’t a clue of the mechanism.
Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
#140This assessment fits with my anecdotal evidence. LLMs just cannot reason in any basic way. LLMs have a large knowledge base that can be spit out at a moment notice. But they have zero insight on its contents, even when the information has just been asked a few lines before. Most of the "intelligence" that LLMs show is just the ability to ask in the correct way the correct questions mirrored back to the user. That is…
LLM reasoning is brittle and not like human cognition, but it is far from zero. It has demonstrably improved to a point where it can solve complex, multi-step problems across domains. See the numerous successful benchmarks and out of sample evals (livebench.ai, imo 2025, trackingai.ai IQ, matharena.ai etc).
I gained multiple months of productivity from vibe coding personally in 2025. If being able to correctly code a complex piece of software from a vague, single paragraph description isn't reasoning, what else is? Btw, I don't code UIs. I code complex mathematical algorithms, some of which never found in textbooks.
> LLMs have a large knowledge base that can be spit out at a moment notice. But they have zero insight on its contents, even when the information has just been asked a few lines before.
LLMs have excellent recall of recent information within their context window. While they lack human-like consciousness or "insight," their ability to synthesize and re-contextualize information from their vast knowledge base is a powerful capability that goes beyond simple data retrieval.
If anything LLMs show polymath-level ability to synthesize information across domains. How do I know? I use them everyday and get great mileage. It's very obvious.
> Most of the "intelligence" that LLMs show is just the ability to ask in the correct way the correct questions mirrored back to the user. That is why there is so many advice on how to do "proper prompting".
Prompting is the user interface for steering the model's intelligence. However, the model's ability to generate complex, novel, and functional outputs that far exceed the complexity of the input prompt shows that its "intelligence" is more than just a reflection of the user's query.
To summarize, I'm appalled by your statements, as a heavy user of SoTA LLMs on a daily basis for practically anything. I suspect you don't use them enough, and lack a viceral feel or scope for their capabilities.