The real hidden message is not that bigger compute produces better results, but that the average user probably doesn't need the top results. In the same way that medium range laptops are now 'good enough' for most people's needs, medium range (e.g. DeepSeek R1x) AI will probably be good enough for most business and user needs. Up till now everyone assumed that only giga-sized server farms could produce anything decen…
Except R1 isn't "medium range" - it's fully competitive with SOTA models at a fraction of the cost. Unless you need multimodal capability or you're desperate to wring out the last percentage point of performance, there's no good reason to use a more expensive model.
The real hidden message is that we're still barely getting started. DeepSeek have completely exploded the idea that LLM architecture has peaked and we're just in an arms race for more compute. 100 engineers found an order of magnitude's worth of low-hanging fruit. What will other companies will be able to do with a similar architectural approach? What other straightforward optimisations are just waiting to be implemented? What will R2 look like if they decide to spend $60m or $600m on a training run?