Great article. > Now, you still want to train the best model you can by cleverly leveraging as much compute as you can and as many trillion tokens of high quality training data as possible, but that's just the beginning of the story in this new world; now, you could easily use incredibly huge amounts of compute just to do inference from these models at a very high level of confidence or when trying to solve extremely…
The impact of competition and DeepSeek on Nvidia
241–250 of 500 posts
Re: The impact of competition and DeepSeek on Nvidia
#242DeepSeek just further reinforces the idea that there is a first-move disadvantage in developing AI models. When someone can replicate your model for 5% of the cost in 2 years, I can only see 2 rational decisions: 1) Start focusing on cost efficiency today to reduce the advantage of the second mover (i.e. trade growth for profitability) 2) Figure out how to build a real competitive moat through one or more of the foll…
This is wrong. First mover advantage is strong. This is why OpenAI is much bigger than Mixtral despite what you said. First mover advantage acquired and keeps subscribers. No one really cares if you matched GPT4o one year later. OpenAI has had a full year to optimize the model, build tools around the model, and used the model to generate better data for their next generation foundational model.
Re: The impact of competition and DeepSeek on Nvidia
#243Great article but it seems to have a fatal flaw. As pointed out in the article, Nvidia has several advantages including: - Better Linux drivers than AMD - CUDA - pytorch is optimized for Nvidia - High-speed interconnect Each of the advantages is under attack: - George Hotz is making better drivers for AMD - MLX, Triton, JAX: Higher level abstractions that compile down to CUDA - Cerbras and Groq solve the interconnect…
In which way? As a user who switched from an AMD-GPU to Nvidia-GPU, I can only report a continued amount of problems with NVIDIAs proprietary driver, and none with AMD. Is this maybe about the open source-drivers or usage for AI?
Re: The impact of competition and DeepSeek on Nvidia
#244Earlier quoted context omitted.
Another aspect that reinforces your point is that the ATM push (and subsequent downfall) was not just bandwidth-motivated but also motivated by a belief that ATM's QoS guarantees were necessary. But it turned out that software improvements, notably MPLS to handle QoS, were all that was needed.
Nah, it's mostly just buffering :-) Plus the cell phone industry paved the way for VOIP by getting everyone used to really, really crappy voice quality. Generations of Bell Labs and Bellcore engineers would rather have resigned than be subjected to what's considered acceptable voice quality nowadays...
Re: The impact of competition and DeepSeek on Nvidia
#245Earlier quoted context omitted.
The difference is the AI researchers have clear plots showing capabilities scaling with GPUs and there's not a sign that it is flattening so they actually have a case for saying that AGI is possible at N GPUs.
Sauce? How do you even measure "capabilities" in that regard, just writing answers to standard tests? Because being able to ace a test doesn't mean it's AGI, it means its good at taking standard tests.
Re: The impact of competition and DeepSeek on Nvidia
#246DeepSeek just further reinforces the idea that there is a first-move disadvantage in developing AI models. When someone can replicate your model for 5% of the cost in 2 years, I can only see 2 rational decisions: 1) Start focusing on cost efficiency today to reduce the advantage of the second mover (i.e. trade growth for profitability) 2) Figure out how to build a real competitive moat through one or more of the foll…
This is wrong. First mover advantage is strong. This is why OpenAI is much bigger than Mixtral despite what you said. First mover advantage acquired and keeps subscribers. No one really cares if you matched GPT4o one year later. OpenAI has had a full year to optimize the model, build tools around the model, and used the model to generate better data for their next generation foundational model.
Re: The impact of competition and DeepSeek on Nvidia
#247Earlier quoted context omitted.
Deepseek is unique, but the US has consistently underestimated Chinese R&D, which is not a winning strategy in iterated games.
There seem to be a 100 fold uptick in jingoists in the last 3-4 years which makes my head hurt but I think there is no consistent "underestimation" in academic circles? I think I have read articles about the up and coming Chinese STEM for like 20 years.
Re: The impact of competition and DeepSeek on Nvidia
#248Earlier quoted context omitted.
> I don't think this seems particularly bad for ChatGPT. They've built a strong brand. This should just help them reduce - by far - one of their largest expenses. Often expenses like that are keeping your competitors away.
Yes, but it typically doesn't matter if someone can reach parity or even surpass you - they have to surpass you by a step function to take a significant number of your users. This is a step function in terms of efficiency (which presumably will be incorporated into ChatGPT within months), but not in terms of end user experience. It's only slightly better there.
Re: The impact of competition and DeepSeek on Nvidia
#249Great article. > Now, you still want to train the best model you can by cleverly leveraging as much compute as you can and as many trillion tokens of high quality training data as possible, but that's just the beginning of the story in this new world; now, you could easily use incredibly huge amounts of compute just to do inference from these models at a very high level of confidence or when trying to solve extremely…
> If most of NVIDIAs moat is in being able to efficiently interconnect thousands of GPUs nah. it moat is CUDA and millions of devs using CUDA aka the ecosystem
So far it seems that the best investment is in RAM producers. Unlike compute the ram requirements seem to be stubborn.
Re: The impact of competition and DeepSeek on Nvidia
#250DeepSeek just further reinforces the idea that there is a first-move disadvantage in developing AI models. When someone can replicate your model for 5% of the cost in 2 years, I can only see 2 rational decisions: 1) Start focusing on cost efficiency today to reduce the advantage of the second mover (i.e. trade growth for profitability) 2) Figure out how to build a real competitive moat through one or more of the foll…
This is wrong. First mover advantage is strong. This is why OpenAI is much bigger than Mixtral despite what you said. First mover advantage acquired and keeps subscribers. No one really cares if you matched GPT4o one year later. OpenAI has had a full year to optimize the model, build tools around the model, and used the model to generate better data for their next generation foundational model.