> While a human may say “aha” to indicate exactly a sudden internal state change, this interpretation is unwarranted for models which do not have any such internal state, and which on the next forward pass will only differ from the pre-aha pass by the inclusion of that single token in their context. Interpreting the “aha” moment as meaningful exemplifies the long-neglected assumption about long CoT models – the false…
By itself, "aha" carries no insight, but the insight is probably stated immediately after it. In that case the aha is semantically useful, by identifying the insight it is near.
Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)
21–30 of 296 posts
Re: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)
#22Earlier quoted context omitted.
By itself, "aha" carries no insight, but the insight is probably stated immediately after it. In that case the aha is semantically useful, by identifying the insight it is near.
it's a rhetorical heuristic that a writer should know to use when directing a reader to a declarative that they want them to pay attention to, usually because it's a non-obvious or roundabout insight when utilized by AI, it's a probabilistic output and it's variable whether or not that rhetorical trick is useful. it also pushes a non-skeptical reader to focus too much on the following text or even to believe that the…
Re: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)
#23If intermediate tokens are not a faithful representation of the computation, then they are a pretty bad audit artifact too. We probably shouldn't be trying to make the model's internal narration more interpretable., but rather the computation around it more reproducible.
Record the actual inputs, model/version/configuration, tool observations and outputs, then make the execution replayable enough that differences between runs can be isolated.
In other words, don't ask the model to explain what it thought, and instead make the system able to show what actually happened.
Re: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)
#24Re: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)
#25Some modest evidence is my own subjective experience of the many times I've explained why I'm doing something, and it is a true explanation in the sense that it is certainly not a lie, but it is also incomplete and there are entire strands of thought that went into my decision that are not being articulated. Though human speech is not equivalent to an LLM's output since we can trivially think without literally speaking whereas they can not. (No need to nitpick on the definitions there; all I'm observing here is that they are forced to emit an externally-visible artifact whereas I can sit in silence, thinking, with no externally-visible artifact being produced. Not trying to make any grand claims about what is "real" cognition or anything.)
It is conceivable how to create a test of whether the tokens correspond to the "real" thought process, and papers and work on that have been done, such as [1]. It is difficult for me to imagine how to scramble the nominal tokens without also completely trashing any implicit calculations that may be occurring too.
[1]: https://transformer-circuits.pub/2025/attribution-graphs/bio...
Re: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)
#26Strong dislike for papers that tell me what to do in the title, especially when even the paper admits a loose correlation of the intermediate tokens compared to solution correctness. My solutions work and they speak for themselves.
I understand the sentiment, and I also use the "thinking" traces as insight, but wouldn't you want your solutions to be based upon a good understanding? If the correlation is weak, then our solution is also weak.
Re: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)
#27Seems like they are closer to scratch than reasoning... Generating some scratch to draw from helps make it easier to compute the real answer.
That's my personal theory too. The model is stuffing its own context with vaguely related tokens, which helps the attention heads retrieve the right tokens.
Re: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)
#28Re: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)
#29Is anthropomorphizing a real problem? From what I know, none of the serious LLM researchers believe it has anything to do with human reasoning, apart from Anthropic with their click-baity terminology like "LLM biology". It's just a metaphor. "Reasoning tokens" is simpler to say than "learned prompt augmentation tokens". I used to (and still do) anthropomorphize things long before LLMs, and I've seen my colleagues do…
A lot of people are not in on the joke. ELIZA effect and AI psychosis is a thing.
Interacting a lot with LLMs might be damaging to the human psyche even for mentally stable people.
Re: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)
#30Is anthropomorphizing a real problem? From what I know, none of the serious LLM researchers believe it has anything to do with human reasoning, apart from Anthropic with their click-baity terminology like "LLM biology". It's just a metaphor. "Reasoning tokens" is simpler to say than "learned prompt augmentation tokens". I used to (and still do) anthropomorphize things long before LLMs, and I've seen my colleagues do…
For example, if you have a search engine or a complex game, you can't run tests like "for all inputs the results are correct", you're going to be fudging a lot, using randomness, using heuristics, and all that kinda stuff
Just like how mathematics > physics > chemistry > biology > psychology > economics/sociology (Auguste Comte's hierarchy reordered a bit for the modern day), moving up the abstraction ladder makes things more complex, less legible and less exact.