Live data from Hacker News

The impact of competition and DeepSeek on Nvidia

youtubetranscriptoptimizer.com

241–250 of 500 posts

Re: The impact of competition and DeepSeek on Nvidia

#241

Great article. > Now, you still want to train the best model you can by cleverly leveraging as much compute as you can and as many trillion tokens of high quality training data as possible, but that's just the beginning of the story in this new world; now, you could easily use incredibly huge amounts of compute just to do inference from these models at a very high level of confidence or when trying to solve extremely…

Running a 680-billion parameter frontier model on a few Macs (at 13 tok/s!) is nuts. That'a two years after ChatGPT was released. That rate of progress just blows my mind.

Re: The impact of competition and DeepSeek on Nvidia

#242

DeepSeek just further reinforces the idea that there is a first-move disadvantage in developing AI models. When someone can replicate your model for 5% of the cost in 2 years, I can only see 2 rational decisions: 1) Start focusing on cost efficiency today to reduce the advantage of the second mover (i.e. trade growth for profitability) 2) Figure out how to build a real competitive moat through one or more of the foll…

This is wrong. First mover advantage is strong. This is why OpenAI is much bigger than Mixtral despite what you said. First mover advantage acquired and keeps subscribers. No one really cares if you matched GPT4o one year later. OpenAI has had a full year to optimize the model, build tools around the model, and used the model to generate better data for their next generation foundational model.

They also burnt a hell of a lot more cash. That’s a disadvantage.

Re: The impact of competition and DeepSeek on Nvidia

#243

Great article but it seems to have a fatal flaw. As pointed out in the article, Nvidia has several advantages including: - Better Linux drivers than AMD - CUDA - pytorch is optimized for Nvidia - High-speed interconnect Each of the advantages is under attack: - George Hotz is making better drivers for AMD - MLX, Triton, JAX: Higher level abstractions that compile down to CUDA - Cerbras and Groq solve the interconnect…

> - Better Linux drivers than AMD

In which way? As a user who switched from an AMD-GPU to Nvidia-GPU, I can only report a continued amount of problems with NVIDIAs proprietary driver, and none with AMD. Is this maybe about the open source-drivers or usage for AI?

Re: The impact of competition and DeepSeek on Nvidia

#244
post #209

Earlier quoted context omitted.

Another aspect that reinforces your point is that the ATM push (and subsequent downfall) was not just bandwidth-motivated but also motivated by a belief that ATM's QoS guarantees were necessary. But it turned out that software improvements, notably MPLS to handle QoS, were all that was needed.

Nah, it's mostly just buffering :-) Plus the cell phone industry paved the way for VOIP by getting everyone used to really, really crappy voice quality. Generations of Bell Labs and Bellcore engineers would rather have resigned than be subjected to what's considered acceptable voice quality nowadays...

Yes, I think most video on the Internet is HLS and similar approaches which are about as far from the ATM circuit-switching approach as it gets. For those unfamiliar HLS is pretty much breaking the video into chunks to download over plain HTTP.

Re: The impact of competition and DeepSeek on Nvidia

#245

Earlier quoted context omitted.

The difference is the AI researchers have clear plots showing capabilities scaling with GPUs and there's not a sign that it is flattening so they actually have a case for saying that AGI is possible at N GPUs.

Sauce? How do you even measure "capabilities" in that regard, just writing answers to standard tests? Because being able to ace a test doesn't mean it's AGI, it means its good at taking standard tests.

This is the canonical paper. Nothing I've seen seems to indicate the curves are flattening, you can ask "scaling what" but the trend is clear.

https://arxiv.org/pdf/2001.08361

Re: The impact of competition and DeepSeek on Nvidia

#246

DeepSeek just further reinforces the idea that there is a first-move disadvantage in developing AI models. When someone can replicate your model for 5% of the cost in 2 years, I can only see 2 rational decisions: 1) Start focusing on cost efficiency today to reduce the advantage of the second mover (i.e. trade growth for profitability) 2) Figure out how to build a real competitive moat through one or more of the foll…

This is wrong. First mover advantage is strong. This is why OpenAI is much bigger than Mixtral despite what you said. First mover advantage acquired and keeps subscribers. No one really cares if you matched GPT4o one year later. OpenAI has had a full year to optimize the model, build tools around the model, and used the model to generate better data for their next generation foundational model.

OpenAI does not have a business model that is cashflow positive at this point and/or a product that gives them a significant leg up in the same moat sense Office/Teams might give to Microsoft.

Re: The impact of competition and DeepSeek on Nvidia

#247

Earlier quoted context omitted.

Deepseek is unique, but the US has consistently underestimated Chinese R&D, which is not a winning strategy in iterated games.

There seem to be a 100 fold uptick in jingoists in the last 3-4 years which makes my head hurt but I think there is no consistent "underestimation" in academic circles? I think I have read articles about the up and coming Chinese STEM for like 20 years.

Precisely. This is the view from the ivory tower.

Re: The impact of competition and DeepSeek on Nvidia

#248

Earlier quoted context omitted.

> I don't think this seems particularly bad for ChatGPT. They've built a strong brand. This should just help them reduce - by far - one of their largest expenses. Often expenses like that are keeping your competitors away.

Yes, but it typically doesn't matter if someone can reach parity or even surpass you - they have to surpass you by a step function to take a significant number of your users. This is a step function in terms of efficiency (which presumably will be incorporated into ChatGPT within months), but not in terms of end user experience. It's only slightly better there.

One data point but my subscription for chatgpt is cancelled every time. So I made every month decision to resub. And because the cost of switching is essentially zero - the moment a better service is up there I will switch in an instant.

Re: The impact of competition and DeepSeek on Nvidia

#249
post #207

Great article. > Now, you still want to train the best model you can by cleverly leveraging as much compute as you can and as many trillion tokens of high quality training data as possible, but that's just the beginning of the story in this new world; now, you could easily use incredibly huge amounts of compute just to do inference from these models at a very high level of confidence or when trying to solve extremely…

> If most of NVIDIAs moat is in being able to efficiently interconnect thousands of GPUs nah. it moat is CUDA and millions of devs using CUDA aka the ecosystem

And then some chineese startup create an amazing compiler that takes cuda and moves it to X (AMD, Intel, Asic) and we are back at square one.

So far it seems that the best investment is in RAM producers. Unlike compute the ram requirements seem to be stubborn.

Re: The impact of competition and DeepSeek on Nvidia

#250

DeepSeek just further reinforces the idea that there is a first-move disadvantage in developing AI models. When someone can replicate your model for 5% of the cost in 2 years, I can only see 2 rational decisions: 1) Start focusing on cost efficiency today to reduce the advantage of the second mover (i.e. trade growth for profitability) 2) Figure out how to build a real competitive moat through one or more of the foll…

This is wrong. First mover advantage is strong. This is why OpenAI is much bigger than Mixtral despite what you said. First mover advantage acquired and keeps subscribers. No one really cares if you matched GPT4o one year later. OpenAI has had a full year to optimize the model, build tools around the model, and used the model to generate better data for their next generation foundational model.

What is OpenAI's first-mover moat? I switched to Claude with absolutely no friction or moat-jumping.
Post reply on HN