Earlier quoted context omitted.
So endless upside might not actually exist ?
It may or may not. I mean it’d be pretty hard to increase yields of GPU fabs or data center sizes another 100x. There are logistical limitations. Unless some Apollo level mission is created by a superpower, we will hit bottlenecks. Algorithmic innovation is the only long term bet.
The trick is always to offset the cost of inference with larger cost of training. You don't apply the Chinchilla scaling law, that's for academics and people who don't have to pay for inference. You pretrain the model 10x longer (like 1T tokens for LLaMA) to make it the best you can fit into an A100 or 4090 quantised to 4 or 3 bits. So everyone can have AI assistants running on their own toys.