I find it interesting because the DeepSeek stuff, while very cool, doesn't seem invalidate that more compute wouldn't translate to even _higher_ capabilities? It's amazing what they did with a limited budget, but instead of the takeaway being "we don't need that much compute to achieve X", it could also be, "These new results show that we can achieve even 1000*X with our currently planned compute buildout" But perhap…
Probably not. If the price of Nvidia is dropping, it's because investors see a world where Nvidia hardware is less valuable, probably because it will be used less. You can't do the distill/magnify cycle like you do with alphago. LLM models have basically stalled in their base capabilities, pre training is basically over at this point, so the news arms race will be over marginal capability gains and (mostly) making th…
Have you considered that compute might be the reason why LLMs are stalled at the moment?
What made LLMs possible in the first place? Right, compute! Transformer Model is 8 years old, technically GPT4 could have been released 5 years ago. What stopped it? Simple, the compute being way too low.
Nvidia has improved compute by 1000x in the past 8 years but what if training GPT5 takes 6-12 months for 1 run based on what OpenAI tries to do?
What we see right now is that pre-training has reached the limits of Hopper and Big Tech is waiting for Blackwell. Blackwell will easily be 10x faster in cluster training (don't look on chip performance only) and since Big Tech intends to build 10x larger GPU clusters then they will have 100x compute systems.
Let's see then how it turns out.
The limit on training is time. If you want to make something new and improve then you should limit training time because nobody will wait 5-6 months for results anymore.
It was fine for OpenAI years ago to take months to years for new frontier models. But today the expectations are higher.
There is a reason why Blackwell is fully sold out for the year. AI research is totally starved for compute.
The best thing for Nvidia is also that while AI research companies compete with each other, they all try to get Nvidia AI HW.