Earlier quoted context omitted.
You had the right understanding in your first comment, but what was missing was the fine tuning. You are right that there aren't many documents on the web that are structured that way, so the raw model wouldn't be very effective on predicting the next token. But since we know that it will complete a command when structured it cleverly, all we had to do to fine tune it is synthesize (generate) a bazillion examples of…
As an example, if you want to see what these sorts of things look like, Databricks open-sourced an instruction fine-tuning dataset sourced from their employees: https://huggingface.co/datasets/databricks/databricks-dolly-... (disclaimer: I'm at Databricks)
I don't like the bland, watered-down tone of ChatGPT, never put together that it's trained on unopinionated data. Feels like a tragedy of the commons thing, the average (or average publically acceptable) view of a group of people is bound to be boring.