Live data from Hacker News

Reasoning models don't always say what they think

anthropic.com

31–40 of 279 posts

Re: Reasoning models don't always say what they think

#31

Earlier quoted context omitted.

When we get to the point where a LLM can say "oh, I made that mistake because I saw this in my training data, which caused these specific weights to be suboptimal, let me update it", that'll be AGI. But as you say, currently, they have zero "self awareness".

> When we get to the point where a LLM can say "oh, I made that mistake because I saw this in my training data, which caused these specific weights to be suboptimal, let me update it", that'll be AGI. While I believe we are far from AGI, I don't think the standard for AGI is an AI doing things a human absolutely cannot do.

We're far from AI. There is no intelligence. The fact the industry decided to move the goal post and re-brand AI for marketing purposes doesn't mean they had a right to hijack a term that has decades of understood meaning. They're using it to bolster the hype around the work, not because there has been a genuine breakthrough in machine intelligence, because there hasn't been one.

Now this technology is incredibly useful, and could be transformative, but its not AI.

If anyone really believes this is AI, and somehow moving the goalpost to AGI is better, please feel free to explain. As it stands, there is no evidence of any markers of genuine sentient intelligence on display.

Re: Reasoning models don't always say what they think

#32

Earlier quoted context omitted.

When we get to the point where a LLM can say "oh, I made that mistake because I saw this in my training data, which caused these specific weights to be suboptimal, let me update it", that'll be AGI. But as you say, currently, they have zero "self awareness".

That’s holding LLMs to a significantly higher standard than humans. When I realize there’s a flaw in my reasoning I don’t know that it was caused by specific incorrect neuron connections or activation potentials in my brain, I think of the flaw in domain-specific terms using language or something like it. Outputting CoT content, thereby making it part of the context from which future tokens will be generated, is roug…

AI CoT may work the same extremely flawed way that human introspection does, and that’s fine, the reason we may want to hold them to a higher standard is because someone proposed to use CoTs to monitor ethics and alignment.

Re: Reasoning models don't always say what they think

#33

One interesting quirk with Claude is that it has no idea its Chain-of-Thought is visible to users. In one chat, it repeatedly accused me of lying about that. It only conceded after I had it think of a number between one and a million, and successfully 'guessed' it.

Edit: 'wahnfrieden corrected me. I incorrectly posited that CoT was only included in the context window during the reasoning task and later left out entirely. Edited to remove potential misinformation.

No, the CoT is not simply extra context the models are specifically trained to use CoT and that includes treating it as unspoken thought

Re: Reasoning models don't always say what they think

#34
post #21

I was under the impression that CoT works because spitting out more tokens = more context = more compute used to "think." Using CoT as a way for LLMs "show their working" never seemed logical, to me. It's just extra synthetic context.

My understanding of the "purpose" of CoT, is to remove the wild variability yielded by prompt engineering, by "smoothing" out the prompt via the "thinking" output, and using that to give the final answer.

Thus you're more likely to get a standardized answer even if your query was insufficiently/excessively polite.

Re: Reasoning models don't always say what they think

#36

Earlier quoted context omitted.

When we get to the point where a LLM can say "oh, I made that mistake because I saw this in my training data, which caused these specific weights to be suboptimal, let me update it", that'll be AGI. But as you say, currently, they have zero "self awareness".

That’s holding LLMs to a significantly higher standard than humans. When I realize there’s a flaw in my reasoning I don’t know that it was caused by specific incorrect neuron connections or activation potentials in my brain, I think of the flaw in domain-specific terms using language or something like it. Outputting CoT content, thereby making it part of the context from which future tokens will be generated, is roug…

Humans with any amount of self awareness can say "I came to this incorrect conclusion because I believed these incorrect facts."

Re: Reasoning models don't always say what they think

#37
post #9

The fact that it was ever seriously entertained that a "chain of thought" was giving some kind of insight into the internal processes of an LLM bespeaks the lack of rigor in this field. The words that are coming out of the model are generated to optimize for RLHF and closeness to the training data, that's it! They aren't references to internal concepts, the model is not aware that it's doing anything so how could it…

>internal concepts, the model is not aware that it's doing anything so how could it "explain itself" This in a nutshell is why I hate that all this stuff is being labeled as AI. Its advanced machine learning (another term that also feels inaccurate but I concede is at least closer to whats happening conceptually) Really, LLMs and the like still lack any model of intelligence. Its, in the most basic of terms, algorith…

While I agree that LLMs are hardly sapient, it's very hard to make this argument without being able to pinpoint what a model of intelligence actually is.

"Human brains lack any model of intelligence. It's just neurons firing in complicated patterns in response to inputs based on what statistically leads to reproductive success"

Re: Reasoning models don't always say what they think

#38
post #19

Earlier quoted context omitted.

> When we get to the point where a LLM can say "oh, I made that mistake because I saw this in my training data, which caused these specific weights to be suboptimal, let me update it", that'll be AGI. While I believe we are far from AGI, I don't think the standard for AGI is an AI doing things a human absolutely cannot do.

All that was described here is learning from a mistake, which is something I hope all humans are capable of.

No, what was described was specifically reporting to an external party the neural connections involved in the mistake and the source in past training data that caused them, as well as learning from new data.

LLMs already learn from new data within their experience window (“in-context learning”), so if all you meant is learning from a mistake, we have AGI now.

Re: Reasoning models don't always say what they think

#39
Humans also post-rationalize the things their subconscious "gut feeling" came up with.

I have no problem for a system to present a reasonable argument leading to a production/solution, even if that materially was not what happened in the generation process.

I'd go even further and pose that probably requiring the "explanation" to be not just congruent but identical with the production would either lead to incomprehensible justifications or severely limited production systems.

Re: Reasoning models don't always say what they think

#40
post #9

The fact that it was ever seriously entertained that a "chain of thought" was giving some kind of insight into the internal processes of an LLM bespeaks the lack of rigor in this field. The words that are coming out of the model are generated to optimize for RLHF and closeness to the training data, that's it! They aren't references to internal concepts, the model is not aware that it's doing anything so how could it…

>internal concepts, the model is not aware that it's doing anything so how could it "explain itself" This in a nutshell is why I hate that all this stuff is being labeled as AI. Its advanced machine learning (another term that also feels inaccurate but I concede is at least closer to whats happening conceptually) Really, LLMs and the like still lack any model of intelligence. Its, in the most basic of terms, algorith…

You are confusing sentience or consciousness with intelligence.
Post reply on HN