Live data from Hacker News

Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)

arxiv.org

91–100 of 296 posts

Re: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)

#91
post #63
post #28

Is anthropomorphizing a real problem? From what I know, none of the serious LLM researchers believe it has anything to do with human reasoning, apart from Anthropic with their click-baity terminology like "LLM biology". It's just a metaphor. "Reasoning tokens" is simpler to say than "learned prompt augmentation tokens". I used to (and still do) anthropomorphize things long before LLMs, and I've seen my colleagues do…

Yes it is a very serious problem because it confuses a lot of folks with a great deal of power like judges and policymakers. The first book I ever read on ML (late 90s) dedicated the entire first or second chapter exploring the distinctions between artificial and biological neurons, and even talked a bit about the philosophy of modelling. I still remember thinking back then why would the authors spend so many pages o…

To be fair, the ANN architecture underneath is a misleading thing to be looking at, it's not where the comparison comes from. Though I can't tell if you meant it to be relevant in that way, or just as a general example for the dangerous nature of metaphor.

LLMs are expressly designed to approximate human behavior within the bounds of the written word. The anthropomorphization is no more philosophically problematic than saying differential calculus measures curves.

Re: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)

#92

Although I 100% agree that the core mechanism of GRPO is purely mechanical token-by-token probability generation, because RL only rewards exact final answers, the training forces the model to develop error-correction habits. This makes the output extremely like human thinking when solving a problem. It's like the order of the thinking tokens is what causes it to get that sweet, delicious reward, and this order seems…

“ This makes the output extremely like human thinking when solving a problem.”

This sounds a little like someone saying a lightbulbs output is extremely like the output of stellar fusion. In one sense, yes. Bulbs are in fact designed to take over when our nearest star is beyond the horizon.

But that really doesn’t mean you call the bulbs mini stars.

Re: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)

#93
post #54

Earlier quoted context omitted.

This is an extremely common fallacy I've seen lots and lots of people fall into with respect to LLMs. In nearly every case, they use the fact that "some humans can't do X" to claim that LLMs are, in fact, basically conscious/human-like/AGI already. This is deeply untrue, and is highly likely to lead them to bad conclusions about what we can and should do with LLMs.

> This is an extremely common fallacy ... they use the fact that "some humans can't do X" to claim that LLMs are, in fact, basically conscious/human-like/AGI already. Claiming that LLMs are conscious or human-like because humans can't do X seems a very strange way to argue for LLM intelligence. Usually, it goes the other way around: an LLM sceptic says "LLMs are dumb because they can't do X" and soon someone has to r…

“ Usually, it goes the other way around: an LLM sceptic says "LLMs are dumb because they can't do X" and soon someone has to remind them that also most of the population can't, in fact, do X.”

And it should not go this way, is the OPs point.

For an analogy, imagine someone looked at a bunch of lightbulbs and expresses dissatisfaction that they aren’t really stars. If someone replies by saying “not all stars are equally bright”, do you think that fact should carry any weight in the argument?

Re: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)

#94
post #86

Earlier quoted context omitted.

I would argue, that the null hypothesis is that it is not, and that anyone claiming that there is a mote of consciousness are the ones with the burden of proof.

The null hypothesis is that we don't know jack shit about consciousness. Any claim of certainty seems extraordinary to me and I want to hear the evidence.

We know quite a lot about consciousness. You may not, but we know enough to know the informational dynamics in a brain are vastly different from an LLMs.

Exactly what kind of certainty are you looking for? Happy to provide it at a molecular, cellular, tissue or whole brain level.

Re: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)

#95

Earlier quoted context omitted.

> This is an extremely common fallacy ... they use the fact that "some humans can't do X" to claim that LLMs are, in fact, basically conscious/human-like/AGI already. Claiming that LLMs are conscious or human-like because humans can't do X seems a very strange way to argue for LLM intelligence. Usually, it goes the other way around: an LLM sceptic says "LLMs are dumb because they can't do X" and soon someone has to r…

“ Usually, it goes the other way around: an LLM sceptic says "LLMs are dumb because they can't do X" and soon someone has to remind them that also most of the population can't, in fact, do X.” And it should not go this way, is the OPs point. For an analogy, imagine someone looked at a bunch of lightbulbs and expresses dissatisfaction that they aren’t really stars. If someone replies by saying “not all stars are equal…

> And it should not go this way, is the OPs point

Well, he's wrong. If you argue that LLMs and humans are fundamentally different because all LLMs do X and no human does it, then showing you that it's not true demolishes your argument. Doesn't prove anything positive, but it certainly proves that your argument is invalid.

Re: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)

#96

Earlier quoted context omitted.

Because plenty of people, even ones that should know better, really believe it's a conscious, thinking entity, not just some turn of phrase. I have a coworker that spends at least 10 hours a week arguing with his like you would with a conscious person. I've gently tried to explain it's like arguing with your compiler for giving you an incoherent error message - it's pointless. It doesn't understand, it can't understa…

Do you have a working definition of "conscious"?

We do have entry and exit conditions. We can disrupt it and study the dynamics of the disruption. We have a good sense of the many mechanisms at play that undergird it.

What we lack is a full map of the exact process from sensation to consciousness across all modalities. And I’m afraid when it comes to interception, we can’t unless we observe every one of the 36 trillion or so cells in a human body, as well as the 36 trillion or so symbiotic and commensal microbes, continuously, all the time.

But I keep finding it astonishing that the claim that we know nothing about consciousness gets bandied about. We know a lot. We don’t have a grand unified theory. The lot we know is definitely split across many levels of evidence and hard to follow, let alone arrange. But this isn’t a black box. It’s a grey box, meeting an even more transparent box that is the LLM, where we do know what the guts are made of, and can interfere at every step in the chain of steps that constitute their dynamics.

Comparing the two, we know there’s a level of similarity in that information gets broken down via a neural network. That similarity was sought.

Since then though, neuroscientists have gone and shown that: 1. The other half of the cells in the brain, the glia, are at least as important as the neurons in cognition and consciousness 2. That interoceptive feedback and feelings are critical drivers of conscious experience 3. Evolutionarily, we know all cells can “learn”, and well before there were neurons or glia or brains, every cell evolved an internal clock that allows it to entrain to external solar and (depending on the species) lunar rhythms. 4. In the last few decades we’ve seen how synaptic activity is shaped and driven both by astrocytes and the circadian clock.

All this is showing us that the abstraction from the 1950s that current neural networks are built on were incomplete.

Whatever these components to do give rise to consciousness in biology (and we’re a long way from done solving this), we certainly wouldn’t imagine with all these modules and mechanisms missing, just maxing on one type of information flow in the brain would give you consciousness.

I’d urge you to not keep insisting consciousness is a total mystery. It’s not, and even your AI model of choice will be able to point you to all the mechanistic evidence we have that whatever it is, it isn’t just neural nets.

Re: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)

#97
post #74
post #58

Earlier quoted context omitted.

> Is anthropomorphizing a real problem? The paper argues that pretending that the so-called thinking traces represent real reasoning can lead users into trusting wrong answers, if the thinking traces appear convincing enough. Researchers might inspect these traces to try to determine the “intent” of a model, as well. For an example of the latter, when OpenAI spoke about the hacking of HuggingFace at Black Hat, they r…

But how is that any different than people being misled by real humans saying words that reflect real thinking, but which are actually dead wrong? The fallacy here is "thinking == correct", not "tokens == thinking"

Because the real thinking still cost the other human the same-ish energy it costs you to put words together, and because after all, the source is a human and not a machine, no, this is very different.

Being mislead may be the shared outcome. But why is different category of source of the mistake and the cost to producer of making the mistake not relevant in this discussion?

Where else in science do you brush aside all differences this way?

Re: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)

#98

Earlier quoted context omitted.

“ Usually, it goes the other way around: an LLM sceptic says "LLMs are dumb because they can't do X" and soon someone has to remind them that also most of the population can't, in fact, do X.” And it should not go this way, is the OPs point. For an analogy, imagine someone looked at a bunch of lightbulbs and expresses dissatisfaction that they aren’t really stars. If someone replies by saying “not all stars are equal…

> And it should not go this way, is the OPs point Well, he's wrong. If you argue that LLMs and humans are fundamentally different because all LLMs do X and no human does it, then showing you that it's not true demolishes your argument. Doesn't prove anything positive, but it certainly proves that your argument is invalid.

Well this requires you to buy the very bullshit argument that comparisons of two physically distinct systems just because they share outputs is meaningful.

If I call a lightbulb an artificial star, the onus is on me to show the behavior under the hood is star like, not just to point at the light and say “you must see it’s a a star since it’s emitting light!”.

Re: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)

#99

Although I 100% agree that the core mechanism of GRPO is purely mechanical token-by-token probability generation, because RL only rewards exact final answers, the training forces the model to develop error-correction habits. This makes the output extremely like human thinking when solving a problem. It's like the order of the thinking tokens is what causes it to get that sweet, delicious reward, and this order seems…

“ This makes the output extremely like human thinking when solving a problem.” This sounds a little like someone saying a lightbulbs output is extremely like the output of stellar fusion. In one sense, yes. Bulbs are in fact designed to take over when our nearest star is beyond the horizon. But that really doesn’t mean you call the bulbs mini stars.

How much is the process of a human child in grade school working through a 3-digit x 3-digit multiplication problem (123 * 456) like a GRPO model with thinking tokens doing the same?

Humans are not born being able to achieve that. It is learned behavior. You and everyone else will remember their teacher saying, "Check your work!" Both the human child and the model work through multiplication problems using the same technique, using the distributive property. They both try to get a reward. For the human child, it is a sense of someone commending them for correctly solving the problem, a reward that probably yields some type of positive dopamine or serotonin feedback loop.

The model solving the problem will have a lower error rate if the first series of tokens created is followed by a series of validation tokens that are subsequently followed by error-correction tokens if there is an error!!!

Maybe it is thinking. Maybe it is remembering to validate and check the work and then remembering to fix the error. For the model trained with RL, why did tokens associated with validation towards the middle of a stream of tokens yield much better results? DeepSeek proved with R1-Zero that a model will learn to verify and correct itself from RL alone with no supervised fine tuning (SFT) teacher ever showing it how. The only reason DeepSeek used SFT was to clean up the reasoning tokens to be human readable. [0] When constrained by SFT, the models will use the double meaning of words -- polysemy -- to satisfy being human-readable while also carrying meaning for what they are working on.

Different people think differently. I watched a viral video of some ~11-year-old child talking to his mom or dad about a stream of a voice in his head. He discovered for the first time that he has a stream of thought. When he goes to school and solves a long multiplication problem, like the stream of tokens from the model, that voice will say to itself (him), "Check your work!"

That is a case of the stream of thought as words being aware of the stream of thoughts as words. Self awareness is a different conversation.

What I think is happening is that the child's stream of thought while solving a multiplication problem in school is likely very similar to an AI model's stream of tokens solving a multiplication problem. And they both were learned. The mechanics are very different, yet, the analogy is apt.

[0] https://huggingface.co/chutesai/DeepSeek-R1-NextN/blob/main/...

Re: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)

#100

Earlier quoted context omitted.

> if it could, you arguing with it isn't going to make it "learn" or act differently. Are you talking about a specific harness that doesn't have context retention mechanisms? For example, ChatGPT with disabled memory feature? Or in general where "it" is a fixed-weights network? The latter is trivially true, of course.

Even claude with “memory” enabled isn’t really “remembering” anything. It just injects it into the context and you hope it happens to find it relevant in its attention mechanisms, and then remembers to actually act on it. Anthropic’s own documentation states claude can and will ignore/truncate these. It’s a context trick, nothing approaching actual “memory,” and in fact, arguing with it will make a bunch of memory fi…

I'm not a big fan of arguments like "it's not the real [human quality], it's [mechanistic explanation]." They lack a part: "because the [human quality] allows us to do X, Y, Z, which is impossible with [this mechanism]."

I agree that the relevance of retrieved pieces and the management of long-term storage could be improved, though.

Post reply on HN