Also see A Better Lesson: https://rodneybrooks.com/a-better-lesson/
Brook's post goes over the classics (Moore's law is ending, curating a dataset requires human intervention, etc) and posits that making a huge model won't be a competitive strategy for long because it gets to expensive to train and use.
It's a bit early to tell, but so far that hasn't materialized. OpenAI got state-of-the-art results with GPT-4, AFAIK by sticking for very-super-big models together. Open source experiments with LLAMA show you can still get good results with heavy quantization. Distillation hasn't be too explored by mainstream projects, but I bet there's lots of potential there too.
Right now the winning strategy looks to be "go really big, then figure out how to go small".