Live data from Hacker News

Nvidia’s $589B DeepSeek rout

finance.yahoo.com

751–760 of 1001 posts

Re: Nvidia’s $589B DeepSeek rout

#751
post #54

I find it interesting because the DeepSeek stuff, while very cool, doesn't seem invalidate that more compute wouldn't translate to even _higher_ capabilities? It's amazing what they did with a limited budget, but instead of the takeaway being "we don't need that much compute to achieve X", it could also be, "These new results show that we can achieve even 1000*X with our currently planned compute buildout" But perhap…

Probably not. If the price of Nvidia is dropping, it's because investors see a world where Nvidia hardware is less valuable, probably because it will be used less. You can't do the distill/magnify cycle like you do with alphago. LLM models have basically stalled in their base capabilities, pre training is basically over at this point, so the news arms race will be over marginal capability gains and (mostly) making th…

Your implication is that we have unlimited compute and therefore know that LLMs are stalled.

Have you considered that compute might be the reason why LLMs are stalled at the moment?

What made LLMs possible in the first place? Right, compute! Transformer Model is 8 years old, technically GPT4 could have been released 5 years ago. What stopped it? Simple, the compute being way too low.

Nvidia has improved compute by 1000x in the past 8 years but what if training GPT5 takes 6-12 months for 1 run based on what OpenAI tries to do?

What we see right now is that pre-training has reached the limits of Hopper and Big Tech is waiting for Blackwell. Blackwell will easily be 10x faster in cluster training (don't look on chip performance only) and since Big Tech intends to build 10x larger GPU clusters then they will have 100x compute systems.

Let's see then how it turns out.

The limit on training is time. If you want to make something new and improve then you should limit training time because nobody will wait 5-6 months for results anymore.

It was fine for OpenAI years ago to take months to years for new frontier models. But today the expectations are higher.

There is a reason why Blackwell is fully sold out for the year. AI research is totally starved for compute.

The best thing for Nvidia is also that while AI research companies compete with each other, they all try to get Nvidia AI HW.

Re: Nvidia’s $589B DeepSeek rout

#752

90% of the comments in this thread make it clear that knowing about technology does not in any way qualify someone to think correctly about markets and equity valuations.

I think that’s correct, but it would be a more useful comment if you gave a few examples and explained the correct thinking.

Re: Nvidia’s $589B DeepSeek rout

#753

90% of the comments in this thread make it clear that knowing about technology does not in any way qualify someone to think correctly about markets and equity valuations.

I think you’re wrong and Wallstreet got Deepseek’s impact wrong. You say DeepSeek should decrease Nvidia demand. Wallstreet agreed today. I say DeepSeek should increase Nvidia’s demand due to Jevon’s Paradox.

> DeepSeek should increase Nvidia’s demand due to Jevon’s Paradox.

How exactly? From what I’ve read the full model can run on MacBook M1 sort of hardware just fine. And this is their first release, I’d expect it to get more efficient and maybe domain specific models can be run on much lower grade hardware sort of raspberry pi sort.

Re: Nvidia’s $589B DeepSeek rout

#755

Earlier quoted context omitted.

I think you’re wrong and Wallstreet got Deepseek’s impact wrong. You say DeepSeek should decrease Nvidia demand. Wallstreet agreed today. I say DeepSeek should increase Nvidia’s demand due to Jevon’s Paradox.

Jevon's paradox ony applies under certain conditions. It remains to be seen if it will hold in this case.

What do you understand those conditions to be?

Re: Nvidia’s $589B DeepSeek rout

#756

Earlier quoted context omitted.

I think you’re wrong and Wallstreet got Deepseek’s impact wrong. You say DeepSeek should decrease Nvidia demand. Wallstreet agreed today. I say DeepSeek should increase Nvidia’s demand due to Jevon’s Paradox.

> DeepSeek should increase Nvidia’s demand due to Jevon’s Paradox. How exactly? From what I’ve read the full model can run on MacBook M1 sort of hardware just fine. And this is their first release, I’d expect it to get more efficient and maybe domain specific models can be run on much lower grade hardware sort of raspberry pi sort.

>How exactly? From what I’ve read the full model can run on MacBook M1 sort of hardware just fine.

No you can't. Unless you run their 1.5b model quantized. Almost useless.

Their full model, Deepseek V3 has 671B parameters. Not even remotely close to being able to run well on consumer hardware.

Re: Nvidia’s $589B DeepSeek rout

#757
post #431

NVIDIA sells shovels to the gold rush. One miner (Liang Wenfeng), who has previously purchased at least 10,000 A100 shovels... has a "side project" where they figured out how to dig really well with a shovel and shared their secrets. The gold rush, wether real or a bubble is still there! NVIDA will still sell every shovel they can manufacture, as soon as it is available in inventory. Fortune 100 companies will still…

AGI would be El Dorado in this analogy?

El Dorado will be AGI with Humanoid robots and actual real pick-axes

Re: Nvidia’s $589B DeepSeek rout

#758

Earlier quoted context omitted.

I believe you that it had to do with the selloff, but I believe that efficiency improvements are good news for NVIDIA: each card just got 20x more useful

That still means that that AI firms don't have to buy as many of Nvidia's chips, which is the whole thing that Nvidia's price was predicated on. FB, Google and Microsoft just had their their billions of dollars in Nvidia GPU capex blown out by $5M side-project. Tech firms are probably not going to be as generous shelling out whatever overinflated price Nvidia was asking for as they were a week ago.

The $5M was the cost of the training itself.

You can rent 10k H100 for 20 days with that money. Go and knock yourself out because that compute is probably higher than what DeepSeek received for that money. And that is public cloud pricing for single H100. I'm sure if you ask for 10k H100 you'll get them at half price so easily 40 days of training.

DeepSeek has fooled everyone by telling them that they need only so less money and people think that they only need to "buy" $5M worth of GPU but that's wrong. The money is the training costs of renting the GPU training hours.

Somebody had to install the 10k GPUs and that's paying $300M to Nvidia.

Re: Nvidia’s $589B DeepSeek rout

#759

Earlier quoted context omitted.

> If training and inference just got 40x more efficient Did training and inference just get 40x more efficient, or just training? They trained a model with impressive outputs on a limited number of GPUs, but DeepSeek is still a big model that requires a lot of resources to run. Moreover, which costs more, training a model once or using it for inference across a hundred million people multiple times a day for a year?…

> DeepSeek is still a big model that requires a lot of resources to run I can run the largest model at 4 tokens per second on a 64GB card. Smaller models are _faster_ than Phi-4. I've just switched to it for my local inference.

Isn't the largest model still like 130GB after heavy quantization[1] and 4 tok/s borderline unusable for interactive sessions with those long outputs?

[1] https://unsloth.ai/blog/deepseekr1-dynamic

Re: Nvidia’s $589B DeepSeek rout

#760

Fascinating. I think Meta is a big winner from this - they still control the content and now mining it for value has been proven even cheaper by DeepSeek.

What kind of quality of content can Meta theoretically mine with all the vast data that they collect? They have a good advertising and a content algorithms but does it approximate general intelligence? And the content you see on Facebook, Instagram, Whatsapp, that's just a lot of average content that might actually be great for AGI.
Post reply on HN