> We (the authors of this website) have at times sought insight into the inner workings of an LLM by asking it “why did you just do that?” > But the LLM can’t tell us. It’s not a person. It doesn’t have the metacognitive abilities necessary to reflect on its past actions and report the motivations underlying them*. > With no clue why it did whatever it just did, the LLM is forced to guess wildly at a plausible explan…
I think this is one of those "some truth, but not the whole truth" things. Yes, we trick ourselves, such as with mis-remembered reflex actions: "I felt it, it hurt, therefore I decided to move", even though the nerve-impulse speeds means your limb was moving before your brain even knew about the pain.
But the "narration" seems to be very important. We create stories to capture cause-and-effect about the world (unclear how much that requires language) and it seems to be beneficially adaptive. In fact this drive is so important that we do it even when we abstractly know it's wrong, like when flipping 50/50 coins and imagining a particular coin is luckier than another, or that you're on a "hot streak", or "now that other outcome is overdue."