The fact that it was ever seriously entertained that a "chain of thought" was giving some kind of insight into the internal processes of an LLM bespeaks the lack of rigor in this field. The words that are coming out of the model are generated to optimize for RLHF and closeness to the training data, that's it! They aren't references to internal concepts, the model is not aware that it's doing anything so how could it…
This article counters a significant portion of what you put forward.
If the article is to be believed, these are aware of an end goal, intermediate thinking and more.
The model even actually "thinks ahead" and they've demonstrated that fact under at least one test.