You didn't really address what og_kalu brought up.
Which is that, it's possible that the model learns human like thinking, because that's the best way to accurately predict the human response itself.
I generally agree with you, but still, I do think that this is the current question. What it is the model is learning that it then uses for predictions?
Because you're assuming it's learning some purely token correlation, like these tokens followed X percent of the time, so that's the response. But it's possible it's learning at a lower level, and understanding the meaning and why those tokens follow in these scenarios, and then applying that same meaning to reasoning process over new tokens in order to predict which one would follow.
I'm skeptical of this, but I do not believe that we really know for sure yet.