Earlier quoted context omitted.
> I mean just try it yourself with o1, go as deep as you like asking how it arrived at a conclusion I don't mean to disagree overall, but on this point the LLM can post-facto rationalize its output but it has no introspection and has absolutely no idea why it made a given bit of output (except in so far as it was a result of COT which it could reiterate to you). The set of weights being activated could be nearly disj…
I doubt people are very accurate at knowing why they made the choices they did. If you want them to recite a chain of reasoning they can but that is kind of far from most decision making most people do.
However we're familiar with the human limits of this and LLMs are currently much worse.
This is particularly relevant because someone suffering from the mistaken belief that LLM's could explain their reasoning might go on to attempt to use that to justify the misapplication of an LLM.
E.g. fine tune some LLM using resume examples so that it almost always rejects Green-skinned people, but approve the LLMs use in hiring decisions because it is insistent that it would never base a decision on someone's skin color. Humans can lie about their biases of course, but a human at least has some experience with themselves while a LLM usually has no experience observing themself except for the output visible in their current window.