Live data from Hacker News

Chain-of-Thought Reasoning in the Wild Is Not Always Faithful (2025)

arxiv.org

41–46 of 46 posts

Re: Chain-of-Thought Reasoning in the Wild Is Not Always Faithful (2025)

#41

Earlier quoted context omitted.

Over time, I've learned to accept that many people -- even very clever ones -- are incapable of holding a metaphor at arm's length. Once they accept the words of a metaphor as applicable at all, the metaphor collapses entirely into literalism for them. They can no longer see that the metaphor was just a tool with inherent limitatation and boundaries. Because the field of artificial "intelligence" is constructed aroun…

I don't think the boundary between "generalization" and "metaphor" is very well defined. When you go from an exemplar of 1 to 2, you're going to find all kinds of edge cases where attributes of the thing being demonstrated that had seemed to be essential turn out to not be necessary. I think you certainly could look at LLMs as "thinking" metaphorically, but I also don't think it is necessarily only a metaphor.

Well, it's unusual to "generalize" an idea if you only had two exemplars, one so old and so complicated that all your terms are specifically referent to it and often even hard to be precise about; and the other is both extremely novel and plainly distinct in both its mechanisms and behaviors.

While maybe that boundary can be fuzzy, we're unequivocally and deeply in "metaphor" territory here.

Re: Chain-of-Thought Reasoning in the Wild Is Not Always Faithful (2025)

#42
post #29

Earlier quoted context omitted.

It’s intentional in the sense that actors are acting intentionally for their own benefit (or at least what they believe is beneficial) and following incentives. Not that there is a master planner who manipulates everything

The current admin seems to have quite a few long term plans they have been working towards. Project 2025, Maralago accords. So far the only major policy item the Trump admin seems to have not intended was the Iran War. Israel killing the intended replacement, Iran leveraging the straight of Hormuz, and dropping three Tomahawks on an elementary school really botched that one.

Yeah, if we are talking about the Trump admin it’s definitely a conspiracy. Pretty much Peter Thiel’s cabal. But they are pretty open about their plans

Re: Chain-of-Thought Reasoning in the Wild Is Not Always Faithful (2025)

#43

Earlier quoted context omitted.

The current political climate is well reasoned and intentional. It might not be yours or mine, however the system is working exactly as the ones paying for it have intended.

My brother and I have been arguing about that all of our lives. He believes everything is intentional and it's just a matter of discovering who benefits. I see chaos that nobody intends or controls. His political landscape is a tapestry of conspiracy theories and mine is a fog of war. I think his is more comforting, since it admits a possibility of a rational, predictable world.

the thing is, that fog might've been true decades ago, but for 100 billionaires to sit in a virtual smokey room and do the shit they want to do, that's not really a conspiracy.

It's just peter theil's texting groups.

The ability to conspiracy both willing and unwilling is such a low threshold now, it's virtually indistinguishable.

You watch one billionaire do something and you're like, I'm a billionaire, I should do that too.

The fact that there's so few billionaires, the probability that they conspire together both direct and indirect approaches 1.

The inverse of course is rediciously hard to conceive: the working class bands together to get something like universal healthcare.

Re: Chain-of-Thought Reasoning in the Wild Is Not Always Faithful (2025)

#44
post #39
post #28

Earlier quoted context omitted.

that sounds testable - if you skip the reasoning tokens, do you get the same result? if not, then there's certainly some bearing, but not necessarily in how we read the tokens as text

It is and has been - skipping the tokens causes performance to drop. But replacing the tokens with filler causes the performance to drop, but by a lot less. You can also train models to emit broken or unrelated thinking tokens and their performance is also not much worse than the ones that are trained to output somewhat coherent thinking traces. This points to a hypothesis that the content of the tokens is only sligh…

I'll never understand some downvoters here.

Re: Chain-of-Thought Reasoning in the Wild Is Not Always Faithful (2025)

#45
post #39

Earlier quoted context omitted.

It is and has been - skipping the tokens causes performance to drop. But replacing the tokens with filler causes the performance to drop, but by a lot less. You can also train models to emit broken or unrelated thinking tokens and their performance is also not much worse than the ones that are trained to output somewhat coherent thinking traces. This points to a hypothesis that the content of the tokens is only sligh…

I'll never understand some downvoters here.

AI boosters denying that their machine god is burbling sweet nothings to itself?

Re: Chain-of-Thought Reasoning in the Wild Is Not Always Faithful (2025)

#46
post #8

Earlier quoted context omitted.

Natural intelligences do this too

Must we always see this restated every time? It's getting a bit stale always seeing these kinds of comments on articles about LLM.

Maybe we can think up a (personal) rule we can follow?

A naked "natural intelligences do this too" might be a bit too short to be useful. But if we can add when/where, cite papers, or show ways in which the parallel operates, then it might be useful.

Compare, eg, talking about a robot arm, and someone goes "a natural arm does this too". You can tell about the fact that it has the same degrees of freedom in the same places, or how this pertains to inverse kinematics, or etc...

Same way here, "this happens to be how natural intelligences seem to solve this too! According to Foo, Bar, Baz et al (2026) the gadget is always twiddled beforehand in macaque apes. " or "Same for natural intelligence: I've noticed I use the same general algorithm myself. I've always considered this the correct way to do translation between languages". --

A more concrete example of a useful answer here.

Natural intelligence does this too! When given the question "explain your reasoning" humans are indeed quite prone to post-hoc confabulation. [1]

[1] https://home.csulb.edu/~cwallis/382/readings/482/nisbett%20s... "Telling More Than We Can Know: Verbal Reports on Mental Processes" (this citation is quite old and may have been superseded, mostly just to illustrate how the rule might work)

Post reply on HN