Earlier quoted context omitted.
Microsoft has been releasing LLMs for years.
Sort of. Phi models were just trained on GPT outputs though.
They also did some more interesting work like showing very small models can be coherent as long as you have very simple children's book style training data (TinyStories is pretty famous).
Lots of these ideas are still used. Learning facts at scale with active reading is an ICLR 2026 paper from Meta AI that does a lot of similar work.