Live data from Hacker News

Nvidia’s $589B DeepSeek rout

finance.yahoo.com

811–820 of 1001 posts

Re: Nvidia’s $589B DeepSeek rout

#811

Earlier quoted context omitted.

> If training and inference just got 40x more efficient Did training and inference just get 40x more efficient, or just training? They trained a model with impressive outputs on a limited number of GPUs, but DeepSeek is still a big model that requires a lot of resources to run. Moreover, which costs more, training a model once or using it for inference across a hundred million people multiple times a day for a year?…

Actually inference got more efficient as well, thanks to the multi-head latent attention algorithm that compresses the key-value cache to drastically reduce memory usage. https://mlnotes.substack.com/p/the-valleys-going-crazy-how-d...

That's a useful performance improvement but it's incremental progress in line with what new models often improve over their predecessors, not in line with the much more dramatic reduction they've achieved in training cost.

Re: Nvidia’s $589B DeepSeek rout

#812
post #625

Earlier quoted context omitted.

I don’t understand why this is not obvious to many people: tech and stock trading are totally two different things, why on earth a tech expert is expected to know trading at all? Imagining how ridiculous it would be if a computer science graduate will also automatically get a financial degree from college even though no financial class has been taken.

I’ve noticed this phenomenon among IT & tech VC crowd. They will launch pod cast, offer expert opinion and what not on just about every topic under the Sun, from cold fusion to COVID vaccine to Ukraine war. You wouldn’t see this in other folks, for example, a successful medical surgeon won’t offer much assertion about NVIDIA. And the general tendency among audience is to assume that expertise can be carried across do…

You totally see this in other folks. Look at Ben Carson, and his magical surgery expertise on running the department of Housing and Urban Development.

Re: Nvidia’s $589B DeepSeek rout

#813
post #807

Earlier quoted context omitted.

It has been clear for a while that one of two things is true. 1) AI stuff isn't really worth trillions, in which case Nvidia is overvalued. 2) AI stuff is really worth trillions, in which case there will be no moat, because you can cross any moat for that amount of money, e.g. you could recreate CUDA from scratch for far less than a trillion dollars and in fact Nvidia didn't spend anywhere near that much to create it…

I think the analysis of (2) is too simplistic because it ignores network effects. A community of developers and users around a specific toolset (e.g. CUDA) is hard to just "buy". Imagine trying to build a better programming language than python -- you could do it for a trillion dollars, but good luck getting the world to use it. For a real example, see Meta and Threads, or any other Twitter competitor.

You have a trillion dollars in incentive. You can use it for more than just creating the software, you can offer incentives to use it or directly contribute patches to the tools people are already using so they support your system. Moreover, third parties already have a large motivation to use any viable replacement because they'd avoid the premium Nvidia charges for hardware.

Re: Nvidia’s $589B DeepSeek rout

#814

Earlier quoted context omitted.

> If training and inference just got 40x more efficient Did training and inference just get 40x more efficient, or just training? They trained a model with impressive outputs on a limited number of GPUs, but DeepSeek is still a big model that requires a lot of resources to run. Moreover, which costs more, training a model once or using it for inference across a hundred million people multiple times a day for a year?…

Actually inference got more efficient as well, thanks to the multi-head latent attention algorithm that compresses the key-value cache to drastically reduce memory usage. https://mlnotes.substack.com/p/the-valleys-going-crazy-how-d...

If H800 is a memory-constrained model that NVIDIA built to avoid the Chinese export ban on H100 with equivalent fp8 performance, it makes zero sense to believe Elon Musk, Dario Armodei and Alexandr Wang's claims that DeepSeek smuggled H100s.

The only reason why a team would allocate time on memory optimizations and writing NVPTX code rather than focusing on posttraining is if they severely struggled with memory during training.

I mean, take a look at the numbers:

https://www.fibermall.com/blog/nvidia-ai-chip.htm#A100_vs_A8...

This is a massive trick pulled by Jensen, take the H100 design whose sales are regulated by the government, make it look 40x weaker and call it H800, while conveniently leaving 8-bit computation as fast as H100. Then bring it to China and let companies stockpile without disclosing production or sales numbers, and have no export controls.

Eventually, after 7 months, US govt starts noticing the H800 sales and introduces new export controls, but it's too late. By this point, DeepSeek has started research using fp8. They slowly build bigger and bigger models, work on the bandwidth and memory consumptions, until they make r1 - their reasoning model.

Re: Nvidia’s $589B DeepSeek rout

#815
post #775
post #695

Earlier quoted context omitted.

I've missed the stories on this until now. Is it known (and is there an ELI5) how they were able to do it so much more efficiently?

This article has good background, context, and explanations [1] They skipped CUDA and instead used PTX which is a lower level instruction set where they were able to implement more performant cross-chip comms to make up for the less-performant H800 chips. [1]: https://stratechery.com/2025/deepseek-faq/

So can people do the same in SPIR for OpenCL or amdgcn?

https://en.wikipedia.org/wiki/Standard_Portable_Intermediate...

https://www.khronos.org/spir/

Or even better in the unified language like SYCL?

https://cdrdv2-public.intel.com/786536/Heidelberg_IWOCL__SYC...

Re: Nvidia’s $589B DeepSeek rout

#816
post #748

Earlier quoted context omitted.

> the models are going to get better/smaller/faster overtime reducing our reliance on the GPU Yes, because we've seen that with other software. I no longer want a GPU for my computer because I play games from the 90s and the CPU has grown powerful enough to suffice... except that's not the case at all. Software grew in complexity and quality with available compute resources and we have no reason to think "AI" will be…

interesting take. the aaa games industry is struggling (e.g. look at the profit warnings, share price drops and studio closures) specifically because people are doing that en masse. but those 90s games are not old - retro has become a movement within gaming and there is a whole cottage industry of "indie" games building that aesthetic because it is cheap and fun.

[dead]

Re: Nvidia’s $589B DeepSeek rout

#817
post #804

Earlier quoted context omitted.

>It's because software devs are smart and make a lot of money They just think they're smart BECAUSE they make a lot of money. Just because you can center divs for six figures a year at a F500 doesn't make you smart at everything.

I've never met a fellow software engineer who "centers divs" for 6 figures. But then I work with engineers using FPGAs to trade in the markets with tick to trade times in double digit nanoseconds and processing streams of market data at ~10 million messages per second (80Gbps) The truth is, a lot of P&L in trading these days is a technical feat of mathematics and engineering and not just one of fundamental analysis a…

>I've never met a fellow software engineer who "centers divs" for 6 figures.

But did you ever meet fellow engineers who don't take everything literally?

Re: Nvidia’s $589B DeepSeek rout

#818

Here’s a take I haven’t seen yet: If training and inference just got 40x more efficient, but OpenAI and co. still have the same compute resources, once they’ve baked in all the DeepSeek improvements, we’re about to find out very quickly whether 40x the compute delivers 40x the performance / output quality, or if output quality has ceased to be compute-bound.

> If training and inference just got 40x more efficient Did training and inference just get 40x more efficient, or just training? They trained a model with impressive outputs on a limited number of GPUs, but DeepSeek is still a big model that requires a lot of resources to run. Moreover, which costs more, training a model once or using it for inference across a hundred million people multiple times a day for a year?…

I think what got cheaper are models with up to date information.

Re: Nvidia’s $589B DeepSeek rout

#819
post #625

Earlier quoted context omitted.

I don’t understand why this is not obvious to many people: tech and stock trading are totally two different things, why on earth a tech expert is expected to know trading at all? Imagining how ridiculous it would be if a computer science graduate will also automatically get a financial degree from college even though no financial class has been taken.

I’ve noticed this phenomenon among IT & tech VC crowd. They will launch pod cast, offer expert opinion and what not on just about every topic under the Sun, from cold fusion to COVID vaccine to Ukraine war. You wouldn’t see this in other folks, for example, a successful medical surgeon won’t offer much assertion about NVIDIA. And the general tendency among audience is to assume that expertise can be carried across do…

Systems, it’s all about systems thinking. It is absolutely true that people in tech are often optimistic and/or delusional about the other expertise at their command. But it’s not like the basic assumption here is completely crazy.

Being a surgeon might require thinking about a few interacting systems, but mostly the number and nature of those systems involved stay the same. Talented programmers without even formal training in CS will eat and digest a dozen brand new systems before breakfast, and model interactions mentally with some degree of fidelity before lunch. And then, any formal training in CS kind of makes general systems just another type of object. This is not the same as how a surgeon is going to look at a heart, or even the body as a whole.

Not that this is the only way to acquire skills in systems thinking. But the other paths might require, IDK, a phd in history/geopolitics, or special studies or extensive work experience in physics or math. And not to rule out other kinds of science or engineering experts as systems thinkers, but a surprisingly large subset of them will specialize and so avoid it. By the numbers.. there are probably just more people in software/IT, therefore more of us to look stupid if/when we get stuff wrong.

Obviously general systems expertise can’t automatically make you an expert on particle physics. But honestly it’s a good piece of background for lots of the wicked problems[1], and the wicked problems are what everyone always wants to talk about.

[1] https://en.m.wikipedia.org/wiki/Wicked_problem

Re: Nvidia’s $589B DeepSeek rout

#820
post #781

90% of the comments in this thread make it clear that knowing about technology does not in any way qualify someone to think correctly about markets and equity valuations.

The crash is absolutely rational; the cascading effect highlights the missing moat for companies like OpenAI. Without a moat, no investor will provide these companies with the billions that fueled most of the demand. This demand was essential for NVIDIA to squeeze such companies with incredible profit margins. NVIDIA was overvalued before, and this correction is entirely justified. The larger impact of DeepSeek is mo…

This is partially why Apple is the one that stands to gain more, and it showed. Their "small models, on device" approach can only be perfected with something like DeepSeek, and they're not exposed to NVIDIA pricing, nor have to prove investors that their approach is still valid.
Post reply on HN