Live data from Hacker News

Nvidia’s $589B DeepSeek rout

finance.yahoo.com

761–770 of 1001 posts

Re: Nvidia’s $589B DeepSeek rout

#761

Earlier quoted context omitted.

That still means that that AI firms don't have to buy as many of Nvidia's chips, which is the whole thing that Nvidia's price was predicated on. FB, Google and Microsoft just had their their billions of dollars in Nvidia GPU capex blown out by $5M side-project. Tech firms are probably not going to be as generous shelling out whatever overinflated price Nvidia was asking for as they were a week ago.

> That still means that that AI firms don't have to buy as many of Nvidia's chips Couldn’t you say that about Blackwell as well? Blackwell is 25x more energy-efficient for generative AI tasks and offer up to 2.5x faster AI training performance overall.

And yet, Blackwell is sold out.

What does that tell us?

The industry is compute starved and that makes totally sense.

The tranformer model on which current LLMs are based on are 8 years old. But why took it so much time to get to the LLMs only 2 years ago?

Simple, Nvidia first had to push the compute at scale strongly. Try training GPT4 on Voltas from 2017. Good luck with that!

Current LLMs are possible thanks to the compute Nvidia has provided in the past decade. You could technically use 20 year old CPUs for LLMs but you might need to connect a billion of them.

Re: Nvidia’s $589B DeepSeek rout

#762

Earlier quoted context omitted.

> If training and inference just got 40x more efficient Did training and inference just get 40x more efficient, or just training? They trained a model with impressive outputs on a limited number of GPUs, but DeepSeek is still a big model that requires a lot of resources to run. Moreover, which costs more, training a model once or using it for inference across a hundred million people multiple times a day for a year?…

> DeepSeek is still a big model that requires a lot of resources to run I can run the largest model at 4 tokens per second on a 64GB card. Smaller models are _faster_ than Phi-4. I've just switched to it for my local inference.

[deleted]

Re: Nvidia’s $589B DeepSeek rout

#763

Earlier quoted context omitted.

> If training and inference just got 40x more efficient Did training and inference just get 40x more efficient, or just training? They trained a model with impressive outputs on a limited number of GPUs, but DeepSeek is still a big model that requires a lot of resources to run. Moreover, which costs more, training a model once or using it for inference across a hundred million people multiple times a day for a year?…

> DeepSeek is still a big model that requires a lot of resources to run I can run the largest model at 4 tokens per second on a 64GB card. Smaller models are _faster_ than Phi-4. I've just switched to it for my local inference.

How are you running it, can you be more specific?

Re: Nvidia’s $589B DeepSeek rout

#764

Earlier quoted context omitted.

> DeepSeek is still a big model that requires a lot of resources to run I can run the largest model at 4 tokens per second on a 64GB card. Smaller models are _faster_ than Phi-4. I've just switched to it for my local inference.

Isn't the largest model still like 130GB after heavy quantization[1] and 4 tok/s borderline unusable for interactive sessions with those long outputs? [1] https://unsloth.ai/blog/deepseekr1-dynamic

OP probably means "the largest distilled model"

Re: Nvidia’s $589B DeepSeek rout

#765

Earlier quoted context omitted.

I think you’re wrong and Wallstreet got Deepseek’s impact wrong. You say DeepSeek should decrease Nvidia demand. Wallstreet agreed today. I say DeepSeek should increase Nvidia’s demand due to Jevon’s Paradox.

Will you post on HN in year describing how much money you earned by buying the stock today?

I did buy:-) although the amount bought isn't impressive because I have other temporary cash flow constraints for a few months.

Re: Nvidia’s $589B DeepSeek rout

#766
post #450

Earlier quoted context omitted.

Jevon's paradox would imply that there's good reason to think that demand for shovels will increase. AI doesn't seem to be one of those things where society as a whole will say, "we have enough of that; we don't need any more". (Many individual people are already saying that, but they aren't the people buying the GPUs for this in the first place. Steam engines weren't universally popular either when they were introdu…

Jevons was talking about coal as an input to commercial processes, for which there were other alternatives that competed on price (e.g. manual/animal labour). Whatever the process, it generated a return, it had utility, and it had scale. I argue it doesn't apply to generative AI because its outputs are mostly no good, have no utility, or are good but only in limited commercial contexts. In the first case, a machine t…

I recently had a discussion with a higher ranked executive and his take on AI changed my outlook a bit. For him the value of ChatGPT:tm: wasn't so much the speed up in any particular task (like presentation generation or so). It's a replacement for consultants.

Yes, the value of those only exists mostly if your internal team is too stubborn to change its opinion. But that seems to be the norm. And the value (those) consultants add is not that high in the first place! They don't have the internal knowledge of _why_ things are fucked up _your particular way_ anyways. That part your team has to contribute anyhow. So the value add shrinks to "throw ideas over the wall and see what sticks". And LLMs are excellent at that.

Yes, that doesn't replace a highly technical consultant that does the actual implementation. Yes, that doesn't give you a good solution. But it probably gives you 5 starting points for a solution before you even finish googling which consultancy to pick (and then waiting for approval and hoping for a goodish team). And that's a story that I can map to reality (not that I like this new bit of information about reality..)

If we accept that story about LLM value, then I think NVIDIA is fine. That generated value is far greater than any amount of energy you can burn on inferring prompts and the only effect will be that the compute-for-training to compute-for-inference ratio decreases further

Re: Nvidia’s $589B DeepSeek rout

#767
post #695

Earlier quoted context omitted.

I've missed the stories on this until now. Is it known (and is there an ELI5) how they were able to do it so much more efficiently?

IIUC they released a paper, it's partially algorithmic improvements partially good old low level optimization.

There's something to be said about the idea that instead of just dumping oceans of money into buying Nvidia cards they just...optimized what they had

I'd say the wider industry could learn a thing or two, but as other commentors have joked. The line must go up

Re: Nvidia’s $589B DeepSeek rout

#768

90% of the comments in this thread make it clear that knowing about technology does not in any way qualify someone to think correctly about markets and equity valuations.

I think you’re wrong and Wallstreet got Deepseek’s impact wrong. You say DeepSeek should decrease Nvidia demand. Wallstreet agreed today. I say DeepSeek should increase Nvidia’s demand due to Jevon’s Paradox.

With all the rate limits on Sonnet 3.5, people only want more of these models. We haven't even seen wide adoption yet.

Re: Nvidia’s $589B DeepSeek rout

#769
post #711

Earlier quoted context omitted.

But did it not lead to net less electricity used for lighting?

With more efficiency offset by more light bulbs, electricity used for lighting has been roughly flat since 2010: https://www.iea.org/data-and-statistics/charts/global-electr... (For LLMs I wish that efficiency could lead to less electricity used for chips, but I think the best we can hope for is for electricity use to flatten out.)

we just have to get solar deployed sufficiently so that we no longer think of energy use as a waste of something precious

Re: Nvidia’s $589B DeepSeek rout

#770

Earlier quoted context omitted.

I think you’re wrong and Wallstreet got Deepseek’s impact wrong. You say DeepSeek should decrease Nvidia demand. Wallstreet agreed today. I say DeepSeek should increase Nvidia’s demand due to Jevon’s Paradox.

I’m also baffled by the reaction. Even with the ability to do more with less, the nature of the race still encourages everyone to do more with more.

I was under the impression too that this would bump the retail customers demand for the 50 series given the extra AI and cuda cores, add to that the relatively low cost of the hardware. But I know nothing of the sentiments around wallstreet.

I don't feel like upgrading my 4090 that said. Maybe wallstreet believes that the larger company deals that have driven the price up for so long might slow down?

Or I'm completely wrong on the impact of the hardware upgrades.

Post reply on HN