The impact of competition and DeepSeek on Nvidia
181–190 of 500 posts
Re: The impact of competition and DeepSeek on Nvidia
#182Earlier quoted context omitted.
Conversely, how much larger can you scale if frontier models only currently need 3 consumer computers? Imagine having 300. Could you build even better models? Is DeepSeek the right team to deliver that, or can OpenAI, Meta, HF, etc. adapt? Going to be an interesting few months on the market. I think OpenAI lost a LOT in the board fiasco. I am bullish on HF. I anticipate Meta will lose folks to brain drain in response…
This assumes no (or very small) diminishing returns effect. I don't pretend to know much about the minutiae of LLM training, but it wouldn't surprise me at all if throwing massively more GPUs at this particular training paradigm only produces marginal increases in output quality.
Re: The impact of competition and DeepSeek on Nvidia
#183Earlier quoted context omitted.
> NVIDIAs moat Offtopic, but your comment finally pushed me over the edge to semantic satiation [1] regarding the word "moat". It is incredible how this word turned up a short while ago and now it seems to be a key ingredient of every second comment. [1] https://en.wikipedia.org/wiki/Semantic_satiation
It is incredible how this word turned up a short while ago… I’m sure if I looked, I could find quotes from Warren Buffet (the recognized originator of the term) going back a few decades. But your point stands.
Re: The impact of competition and DeepSeek on Nvidia
#184Great article. > Now, you still want to train the best model you can by cleverly leveraging as much compute as you can and as many trillion tokens of high quality training data as possible, but that's just the beginning of the story in this new world; now, you could easily use incredibly huge amounts of compute just to do inference from these models at a very high level of confidence or when trying to solve extremely…
Re: The impact of competition and DeepSeek on Nvidia
#185 Perhaps most devastating is DeepSeek's recent efficiency breakthrough, achieving comparable model performance at approximately 1/45th the compute cost. This suggests the entire industry has been massively over-provisioning compute resources.
I wrote in another thread why DeepSeek should increase demand for chips, not lower.1. More efficient LLMs should lead to more usage, which means more AI chip demand. Jevon's Paradox.
2. Even if DeepSeek is 45x more efficient (it is not), models will just become 45x+ bigger. It won’t stay small.
3. To build a moat, OpenAI and American AI companies need to up their datacenter spending even more.
4. DeepSeek's breakthrough is in distilling models. You still need a ton of compute to train the foundational model to distill.
5. DeepSeek's conclusion in their paper says more compute is needed for next break through.
6. DeepSeek's model is trained on GPT4o/Sonnet outputs. Again, this reaffirms the fact that in order to take the next step, you need to continue to train better models. Better models will generate better data for next-gen models.
I think DeepSeek hurts OpenAI/Anthropic/Google/Microsoft. I think DeepSeek helps TSMC/Nvidia.
Combined with the emergence of more efficient inference architectures through chain-of-thought models, the aggregate demand for compute could be significantly lower than current projections assume.
This is misguided. Let's think logically about this.More thinking = smarter models
Faster hardware = more thinking
More/newer Nvidia GPUs, better TSMC nodes = faster hardware
Therefore, you can conclude that Nvidia and TSMC demand should go up because of CoT models. In 2025, CoT models are clearly bottlenecked by not having enough compute.
The economics here are compelling: when DeepSeek can match GPT-4 level performance while charging 95% less for API calls, it suggests either NVIDIA's customers are burning cash unnecessarily or margins must come down dramatically.
Or that in order to build a moat, OpenAI/Anthropic/Google and other laps need to double down on even more compute.Re: The impact of competition and DeepSeek on Nvidia
#186Perhaps most devastating is DeepSeek's recent efficiency breakthrough, achieving comparable model performance at approximately 1/45th the compute cost. This suggests the entire industry has been massively over-provisioning compute resources. I wrote in another thread why DeepSeek should increase demand for chips, not lower. 1. More efficient LLMs should lead to more usage, which means more AI chip demand. Jevon's Par…
Re: The impact of competition and DeepSeek on Nvidia
#187Perhaps most devastating is DeepSeek's recent efficiency breakthrough, achieving comparable model performance at approximately 1/45th the compute cost. This suggests the entire industry has been massively over-provisioning compute resources. I wrote in another thread why DeepSeek should increase demand for chips, not lower. 1. More efficient LLMs should lead to more usage, which means more AI chip demand. Jevon's Par…
Fwiw many of the improvements in Deepseek were already in other 'can run on your personal computer' AI's such as Meta's Llama. Deepseek is actually very similar to Llama in efficiency. People were already running that on home computers with M3's.
A couple of examples; Meta's multi-token prediction was specifically implemented as a huge efficiency improvement that was taken up by Deepseek. REcurrent ADaption (READ) was another big win by Meta that Deepseek utilized. Multi-head Latent Attention is another technique, not pioneered by Meta but used by both Deepseek and Llama.
Anyway Deepseek isn't some independent revolution out of nowhere. It's actually very very similar to the existing state of the art and just bundles a whole lot of efficiency gains in one model. There's no secret sauce here. It's much better than what openAI has but that's because openAI seem to have forgotten 'The Bitter Lesson'. They have been going at things in an extremely brute force way.
Anyway why do i point out that Deepseek is very similar to something like Llama? Because Meta's spending 100's of billions on chips to run it. It's pretty damn efficient, especially compared to openAI but they are still spending billions on datacenter build-outs.
Re: The impact of competition and DeepSeek on Nvidia
#188Perhaps most devastating is DeepSeek's recent efficiency breakthrough, achieving comparable model performance at approximately 1/45th the compute cost. This suggests the entire industry has been massively over-provisioning compute resources. I wrote in another thread why DeepSeek should increase demand for chips, not lower. 1. More efficient LLMs should lead to more usage, which means more AI chip demand. Jevon's Par…
I agree with this. Fwiw many of the improvements in Deepseek were already in other 'can run on your personal computer' AI's such as Meta's Llama. Deepseek is actually very similar to Llama in efficiency. People were already running that on home computers with M3's. A couple of examples; Meta's multi-token prediction was specifically implemented as a huge efficiency improvement that was taken up by Deepseek. REcurrent…
Isn't the point of 'The Bitter Lesson' precisely that in the end, brute force wins, and hand-crafted optimizations like the ones you mention llama and deepseek use are bound to lose in the end?
Re: The impact of competition and DeepSeek on Nvidia
#189Re: The impact of competition and DeepSeek on Nvidia
#190Perhaps most devastating is DeepSeek's recent efficiency breakthrough, achieving comparable model performance at approximately 1/45th the compute cost. This suggests the entire industry has been massively over-provisioning compute resources. I wrote in another thread why DeepSeek should increase demand for chips, not lower. 1. More efficient LLMs should lead to more usage, which means more AI chip demand. Jevon's Par…
But Microsoft hosts 3rd party models too, and cheaper models means more usage, which means more $$$ to scaled cloud providers right?