Earlier quoted context omitted.
But why is "competing against remote SOTA models on quality" the only thing that matters here?
What the hell else is there? All the other stuff can be done by an intern with an 8 euro HF Pro subscription. Other than actual research, which is in a different camp.
Besides that, there is a ton of use cases for smaller models for a bunch of different things. We'll be unlikely to be able to run LLMs (actually Large) on smartphones for a while, while the smaller LLMs seem to run already on-device in experiments.