Earlier quoted context omitted.
This is just a baseless conspiracy theory that I have, but I do wonder if they intentionally avoided certain avenues of research because it could erode the “moat” major tech firms have by dramatically reducing the capital costs required for training and inference. If you focus your research around work that mandates high-end hardware at scale you can lock out tons of potential incumbents
There is nothing in the deepseek paper that suggests you can't use the order of magnitude in hardware costs you saved to just train models that are ten times as large.
There's already some pretty impressive work being done with folks using just a pair of M2 Ultras with r1 in a "home lab" context that goes way beyond what you could previously do with llama.