Live data from Hacker News

Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens

arstechnica.com

101–110 of 140 posts

Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens

#101

Earlier quoted context omitted.

100% productivity gains on coding tasks are absolutely within the realm of possibility

If that is the case, I would argue that you were taking money for doing a job that should've been automated or abstracted already.

I would disagree. I would argue that if you aren't seeing gains in your productivity, you're either using the tools incorrectly, or you are in some ultra specific niche area of coding that AI isn't helpful on yet.

Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens

#102

Earlier quoted context omitted.

Not sure what all this is about, I somewhat regret taking a breaking from coding with LLMs to have it explained to me its all a mirage and a secret and sloppy plan for getting me an automagic egirl or something. ;)

Right? Oh this fairly novel solution the the problem I was having that works and is well tested. Oh throw it away.. sorry the model can't think of stuff.. Back to square one!!

Can you please share a few sessions ? I want to get a better sense of what people have achieved with generic LLMs that is novel. (Emphasis on "generic", I think I can more readily imagine how specialized models for protein folding can lead to innovation)

Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens

#103
post #78

Earlier quoted context omitted.

How about the International Math Olympiad? https://arstechnica.com/ai/2025/07/google-deepmind-earns-gol...

You're saying they don't use math textbooks and math forums to train LLMs, then?

The problems are not in textbooks. I’m curious what would count as an out of distribution problem for you. Only problems no one knows how to solve?

Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens

#104
post #78

Earlier quoted context omitted.

How about the International Math Olympiad? https://arstechnica.com/ai/2025/07/google-deepmind-earns-gol...

You're saying they don't use math textbooks and math forums to train LLMs, then?

You can apply this same argument to humans, 99.999% of people will not be able to escape it.

In the case of the Math Olympiad, the students who take it grind hours a day for months on practice problems and past Olympiad problems.

Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens

#105

It's interesting that there's still such a market for this sort of take. > In a recent pre-print paper, researchers from the University of Arizona summarize this existing work as "suggest[ing] that LLMs are not principled reasoners but rather sophisticated simulators of reasoning-like text." What does this even mean? Let's veto the word "reasoning" here and reflect. The LLM produces a series of outputs. Each output c…

If you read the comments of AI articles on Arstechnica, you will find that they seem to have becomes the tech bastion of anti-ai. I'm not sure how it happened, but it seems they found or fell into a strong anti-AI niche, and now feed it.

You cannot even see the comments of people who pointed out the flaws in the study, since they are so heavily downvoted.

Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens

#106
post #99
post #53

> ... that these "reasoning" models can often produce incoherent, logically unsound answers when questions include irrelevant clauses or deviate even slightly from common templates found in their training data. I have encountered this problem numerous times, now. It really makes me believe that the models do not really understand the topic, even the basics but just try to predict the text. One recent example was me a…

> It really makes me believe that the models do not really understand the topic, even the basics but just try to predict the text. This is correct. There is no understanding, there aren't even concepts. It's just math, it's what we've been doing with words in computers for decades, just faster and faster. They're super useful in some areas, but they're not smart, they don't think.

I’ve never seen so much misinformation trotted out by the laity as I have with LLMs. It’s like I’m in a 19th century forum with people earnestly arguing that cameras can steal your soul. These people haven’t a clue of the mechanism.

Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens

#108
post #76

> LLMs are [...] sophisticated simulators of reasoning-like text Most humans are unsophisticated simulators of reasoning-like text.

Except you, right? You're one of the special few who can actually reason, not like /those/ people.

You're completely missing the point of OP's comment, and strangely, ironically lending credence to your interpretation of that comment lol (self-inflicted harm).

We don't have a good scientific or philosophical handle on what it actually means to "think" (let alone consciousness).

Humanity has so far been really bad at even using relative heuristics based on our own experiences to recognize, classify, and reason about entities that "think."

So it's really amusing when authors just arbitrarily side-step this whole issue and describe these systems as categorically not being real but imitating the real thing... all the while not realizing such characterizations apply to humanity as well.

Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens

#109
post #80
post #67

Earlier quoted context omitted.

This is why research like this is important and needs to keep being published. What we have seen the last few years is a conscious marketing effort to rebrand everything ML as AI and to use terms like "Reasoning", "Extended Thinking" and others that for many non technical people give the impression that it is doing far more than it is actually doing. Many of us here can see his research and be like... well yeah we al…

I'm happy enough if I'm better off for having used a tool than having not.

Most people weren’t happy when the 2008 crash happened, and bank bailouts were needed, and a global recession ensued.

Most people here are going to use a coding agent, be happy about it (like you), and go on their merry way.

Most people here are not making near trillion dollar bets on the world changing power of AI.

EVERYONE here will be affected by those bets. It’s one thing if those bets pay off if future subscription growth matches targets. It’s an entirely different thing if those bets require “reasoning” to pan out.

Re: Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens

#110

> Without specification, we employ a decoder-only language model GPT2 (Radford et al., 2019) with a configuration of 4 layers, 32 hidden dimensions, and 4 attention heads. Yeah, ok. The research is interesting, warranted, but writing an article about it, and leading with the conclusions gathered from toy models and implying this generalises to production LLMs is useless. We've been here before with small models. Trai…

The results from a smaller model are still viable if the paradigm is identical. Unless you believe that larger volumes of data leads to more (unexplained) emergent properties of the AI. i.e, if you think that a larger volume of training data somehow means the model develops actual reasoning skills, beyond the normal next-token prediction.

I do think that larger models will perform better, but not because they fundamentally work differently than the smaller models, and thus the idea behind TFA still stands (in my opinion).

Post reply on HN