Earlier quoted context omitted.
This is adapted from Microsoft research's YOCO. It was known for a while(2024!). Yes, credit to Deepseek for actually scaling it up and releasing a frontier flash LLM. Edit: the rest of this thread has become a US China infowar theory culture war. I am not of either of these countries and the above comment isnt meant to implicitly support either "side".
why didn't Microsoft scale its own invention?
Thus, the reality is that most of these ideas become polished only when its actually deployed and it has to work outside of a PoC. Since LLMs are a high capex product, only very few people actually make non-PoCs. Deepseek is in the business of low cost, fast inference. So they are the ones actually polishing these efficiency-ish ideas and combining many of them (this one, then engram which is based on multiple previous ideas including google brain's ngrammer) to make a coherent system. Openai and anthropic's systems will also involve a polished combination of multiple ideas for each of their systems - Luna is likely a combination of a few efficiency-ish ideas. Shame they won't publish though.
As for microsoft, they don't really sell models, they sell azure. So there is no reason for them to do the high capex scale out of these types of bags of techniques. In a sense, it did benefit them, others developed the model and now many US customers can serve DS4.1 Flash on Azure datacenters.
If it is not clear, I am not understating anything. Combining these rough ideas and making them work actually involves real novel ideas on top and is what is much more difficult than the academic results that were built upon. This also does not mean that the academic results are useless, they are what give us useful priors at all in what is a highly empirical field.