Earlier quoted context omitted.
There is nothing in the deepseek paper that suggests you can't use the order of magnitude in hardware costs you saved to just train models that are ten times as large.
But there are thresholds of commercial viability in all of this. DeepSeek's technology gets you over that line with less sophisticated hardware. There's already some pretty impressive work being done with folks using just a pair of M2 Ultras with r1 in a "home lab" context that goes way beyond what you could previously do with llama.
There are things, like doc strings for functions, that llama models give better results for.