Live data from Hacker News

Nvidia’s $589B DeepSeek rout

finance.yahoo.com

161–170 of 1001 posts

Re: Nvidia’s $589B DeepSeek rout

#161

Good time to buy then, I don’t understand how stupid some traders can be. A more efficient model is better for NVIDIA not worse. More compute is still better given the same model. And as more efficient models proliferate it means more edge computing which means more customers with lower negotiating power than Meta and Google… This is like thinking that if people need to dig only 1 day instead of an entire month to ge…

Indeed: https://en.wikipedia.org/wiki/Jevons_paradox In economics, the Jevons paradox occurs when technological progress increases the efficiency with which a resource is used, but the falling cost of use induces increases in demand enough that resource use is increased, rather than reduced.

Yes but it will also mean that people wouldn't need cutting edge NVIDIA chips - it will be able to run on older node chips. Or from different manufacturers. So NVIDIA wouldn't be able to command the margins they do now.

It may be great news for VRAM manufacturers tough.

Re: Nvidia’s $589B DeepSeek rout

#162
post #2

Nvidia -13% in Frankfurt stock market just now. Valuations of private unicorns like OpenAi and Anthropic must be in free fall. DeepSeek spends $6 million in old H800 hardware to develop open source model that overtakes ChatGPT. AI gets better, but profit margins sink with strong competition. Chinese AI startup DeepSeek overtakes ChatGPT on Apple App Store https://news.ycombinator.com/item?id=42839656 Edit: Nvidia now…

> overtakes ChatGPT

That's arguable, though. I mean it's much cheaper and reasonably competitive which is almost the same but IMHO DeepSeek seems to get stuck in random loops and hallucinates more frequently than o1.

Re: Nvidia’s $589B DeepSeek rout

#163
post #15

Earlier quoted context omitted.

Their model is open and they published paper describing it https://arxiv.org/pdf/2412.19437 The can't be far off, or it would be noticed. Even if they are heavily government subsidized for energy and hardware, I don see how the cost of training in the US would be more than double.

They express their cost in terms of GPU hours, then convert that to USD based on market GPU rental rates, so it's not affected by subsidies. It's possible however they lied about GPU hours, but if that was the case an expert should be able to show they lied by working out how many flops are needed to train based on the amount of tokens they say they used vs the flops of the GPUs they say they used.

Total training FLOPs can be deduced from model architecture (which they can't hide since they released weights) and how many tokens they trained on. With total training FLOPs and GPU hours you can calculate MFU. And the MFU of their deepseek-v3 train is around 40%, which sounds right. Both Google and Meta reported higher MFU. So the GPU hours should be correct. The only thing they could have lied is on how many tokens they trained the model on. DeepSeek reported 14T which is also similar to what Meta did so nothing crazy here.

tl;dr all numbers check up and the winnings come from the model architecture innovations they made.

Re: Nvidia’s $589B DeepSeek rout

#164

Earlier quoted context omitted.

agreed but ASML has been chronically undervalued forever.

Have you priced in the extremely limited freedom to operate they have? There is an extreme systemic risk to being a monopoly in a strategic position. It's an extreme beneficial position to be in, until it isn't.

do you have a model for "short-leash monopolies" would definitely put that in The Book.

Re: Nvidia’s $589B DeepSeek rout

#165

The parallels with the dotcom bubble are clear... Tech is magical. Tech becomes cheap/free. Stock holders panic. For some time it's been clear that there's an AI bubble and this may be what finally pops it.

I don't think there has to be an AI bubble, but valuations overall have to come down to something in accordance with the interest rates and expected long-term profit rates.

Re: Nvidia’s $589B DeepSeek rout

#166

Earlier quoted context omitted.

Meta should actually go up from this. If deep seek is perfect they don't need to pay for an expensive Llama team. And even if deep seek isn't perfect, the low cost training strategies they've invented could be used by Meta to reduce the cost of Llama training. Although Meta develops models they don't sell them. So a world where foundation models are free is fine for them.

But if they don't sell them, are they spending 50 billion a year on open source goodwill?

It's like everyone forgets that App Tracking Transparency (ATT) was supposed to put Meta out of business. By many accounts, Meta's ad targeting is even better now than before ATT. It's been reported that AI is what saved their ad targeting.

The OSS goodwill is just a side effect and a way to undermine companies who are not using AI to effectively make profits today.

Cheaper/more efficient is absolutely great for Meta. If they can lower their capex it would be an instant bump to their bottom line.

Re: Nvidia’s $589B DeepSeek rout

#167
post #2

Nvidia -13% in Frankfurt stock market just now. Valuations of private unicorns like OpenAi and Anthropic must be in free fall. DeepSeek spends $6 million in old H800 hardware to develop open source model that overtakes ChatGPT. AI gets better, but profit margins sink with strong competition. Chinese AI startup DeepSeek overtakes ChatGPT on Apple App Store https://news.ycombinator.com/item?id=42839656 Edit: Nvidia now…

> DeepSeek spends $6 million in old H800 hardware to develop open source model that overtakes ChatGPT. DeepSeek claims that's what they spent. They're under a trade embargo, and if they had access to any more than that it would have been obtained illegally. They might be telling the truth, but let's wait until someone else replicates it before we fully accept it.

>been obtained illegally.

PRC companies breaking US export control laws is legal (for PRC companies). Maybe they're trying to avoid US entity listing, lot's of PRC companies keep mum about growing capabilites to do so. But the mere fact Deepseek is publicizing means they're unlikely to care about the political heat that is coming and the ramifications. If anything, getting on US entity list probably locks in their employees with Deepseek on resume into PRC.

Re: Nvidia’s $589B DeepSeek rout

#168
post #72
post #54

I find it interesting because the DeepSeek stuff, while very cool, doesn't seem invalidate that more compute wouldn't translate to even _higher_ capabilities? It's amazing what they did with a limited budget, but instead of the takeaway being "we don't need that much compute to achieve X", it could also be, "These new results show that we can achieve even 1000*X with our currently planned compute buildout" But perhap…

The stock market is not the economy, Wall Street is not Main Street. You need to look at this more macroscopically if you want to understand this. Basically: China tech sector just made a big splash, traders who witnessed this think other traders will sell because maybe US tech sector wasn't as hot, so they sell as other traders also think that and sell. The fall will come to rest once stocks have fallen enough that…

Is it really related to China's tech sector as such, though? If this is true then Openai, Google or even many magnitudes smaller companies etc. can just easily replicate similar methods in their processes and provide models which are just as good or better. However they'll need way less Nvidia GPUs and other HW to do that than when training their current models.

Re: Nvidia’s $589B DeepSeek rout

#169
post #87

Earlier quoted context omitted.

Andrej Karpathy was tweeting about DeepSeek a month (!) ago. "DeepSeek (Chinese AI co) making it look easy today with an open weights release of a frontier-grade LLM trained on a joke of a budget (2048 GPUs for 2 months, $6M)." https://x.com/karpathy/status/1872362712958906460

this same forum ignored deepseek one month ago, save for a few... open minded people.

"That forum", just like the mobs in ancient Rome, doesn't operate on logic, it operates on emotional feelings and groupthink.

Re: Nvidia’s $589B DeepSeek rout

#170

The part of this that doesn’t jibe with me is the fact that they also released this incredibly detailed technical report on their architecture and training strategy. The paper is well-written and has a lot of specifics. Exactly the opposite of what you would do if you had truly made an advancement of world-altering magnitude. All this says to me is that the models themselves have very little intrinsic value / are hig…

What has changed is the perception that people like OpenAI/MSFT would have an edge on the competition because of their huge datacenters full of NVDA hardware. That is no longer true. People now believe that you can build very capable AI applications for far less money. So the perception is that the big guys no longer have an edge.

Tesla had already proven that to be wrong. Tesla's Hardware 3 is a 6 year old design, and it does amazingly well on less than 300 watts. And that was mostly trained on a 8k cluster.

Post reply on HN