Live data from Hacker News

Nvidia’s $589B DeepSeek rout

finance.yahoo.com

401–410 of 1001 posts

Re: Nvidia’s $589B DeepSeek rout

#401

Earlier quoted context omitted.

there's no reason to believe that performance will continue to scale with compute, though. that's why there's a rout. more simply, if you assume maximum performance with the current LLM/transformer architecture is say, twice as good as what humanity is capable of now, then that would mean that you're approaching 50%+ performance with orders of magnitude less compute. there's just no way you could justify the amount o…

Wait no, there is actually PLENTY of evidence that performance continues to scale with more compute. The entire point of the o3 announcement and benchmark results of throwing a million bucks of test time compute at ARC-AGI is that the ceiling is really really high. We have 3 verified scaling laws of pre-training corpus size, parameter count, and test time compute. More efficiency is fantastic progress, but we will al…

there's evidence that performance increases with compute, but not that it scales with compute, e.g. linearly or exponentially. the SOTA models already are seeing diminishing returns w.r.t parameter size, training time and generally just engineering effort. it's a fact that doubling, say, parameter size does not double benchmark performance.

would love to see evidence to the contrary. my assertion comes from seeing claude, gemini and o1.

if anything I feel performance is more of a function of the quality of data than anything else.

Re: Nvidia’s $589B DeepSeek rout

#402
post #335

Still overvalued IMO. Their market cap remains ludicrous.

Everything is overvalued if you think in terms of earning multiples, p/s, p/e should be 1:1

No, that means that they're earning enough in one year to cover their entire valuation. You want something like 10:1 p/e which means that the next 10 years earnings are factored in to cover their present valuation.

Re: Nvidia’s $589B DeepSeek rout

#403

Earlier quoted context omitted.

Hype buyers are also Hype sellers - anything Nvidia was last week is exactly what it is this week - DeepSeek doesn't really have any impact on Nvidia sales - Some argument could be made that this can shift compute off of cloud and onto end user devices, but that really seems like a stretch given what I've seen running this locally.

I agree hype is a big portion of it, but if DeepSeek really has found a way to train models just as good as frontier ones for a hundredth of the hardware investment, that is a substantial material difference for Nvidia's future earnings.

1. Nobody has replicated their DeepSeek's results on their reported budget yet. Scale.ai's Alexander Wang says they're lying and that they have a huge, clandestine H100 cluster. HuggingFace is assembling an effort to publicly duplicate the paper's claims.

2. Even if DeepSeek's budget claims are true, they trained their model on the outputs of an expensive foundation model built from a massive capital outlay. To truly replicate these results from scratch, it might require an expensive model upstream.

Re: Nvidia’s $589B DeepSeek rout

#404
post #341

IMO this is less about DeepSeek and more that Nvidia is essentially a bubble/meme stock that is divorced from the reality of finance and business. People/institutions who bought on nothing but hype are now panic selling. DeepSeek provided the spark, but that's all that was needed, just like how a vague rumor is enough to cause bank runs.

I don't think it's fair to say NVDA is meme stock, having reported 35B revenue last quarter.

Nvidia's annual revenue in 2024 was $60B. In comparison, Apple made $391B. Microsoft made $245B. Amazon made $575B. Google made $278B. And Nvidia is worth more than all of them. You'd have to go very far down the list to find a company with a comparable ratio of revenue or income to market cap as Nvidia.

Re: Nvidia’s $589B DeepSeek rout

#405
post #370

This is really dumb. Deepseek showing that you can do pure online RL for LLMs means we now have a clear path to just keep throwing more compute at the problem! If anything we made the whole "we are hitting a data wall" problem even smaller. Additionally, its yet another proof point that scaling inference compute is a way forward. Models that think for hours or days are the future. As we move further into the regime o…

Didn't DeepSeek also show that pure RL leads to low-quality results compared to also doing old-fashioned supervised learning on a "problem solving step by step" dataset? I'm not sure why people are getting excited about the pure-RL approach, seems just overly complicated for no real gain.

Re: Nvidia’s $589B DeepSeek rout

#406

Earlier quoted context omitted.

Hype buyers are also Hype sellers - anything Nvidia was last week is exactly what it is this week - DeepSeek doesn't really have any impact on Nvidia sales - Some argument could be made that this can shift compute off of cloud and onto end user devices, but that really seems like a stretch given what I've seen running this locally.

I agree hype is a big portion of it, but if DeepSeek really has found a way to train models just as good as frontier ones for a hundredth of the hardware investment, that is a substantial material difference for Nvidia's future earnings.

Is it? Training is only done once, inference requires GPUs to scale, especially for a 685B model. And now, there’s an open source o1 equivalent model that companies can run locally, which means that there’s a much bigger market for underutilized on-prem GPUs.

Re: Nvidia’s $589B DeepSeek rout

#407
post #341

IMO this is less about DeepSeek and more that Nvidia is essentially a bubble/meme stock that is divorced from the reality of finance and business. People/institutions who bought on nothing but hype are now panic selling. DeepSeek provided the spark, but that's all that was needed, just like how a vague rumor is enough to cause bank runs.

Not quite, I believe this sell off was caused by DeepSeek showing with their new model that the hardware demands of AI are not necessarily as high as everyone has assumed (as required by competing models). I've tried their 7b model, running locally on a 6gb laptop GPU. Its not fast, but the results I've had have rivaled GPT4. Its impressive.

I believe you that it had to do with the selloff, but I believe that efficiency improvements are good news for NVIDIA: each card just got 20x more useful

Re: Nvidia’s $589B DeepSeek rout

#408

So the Chinese graciously gift a paper and model which describes methods that radically increase the efficiency of hardware which will allow US AI firms to create much better models due to having significantly more AI hardware and people are bearish on US AI now?

> the Chinese

Daya Guo, Dejian Yang, Haowei Zhang, et.al., quant researchers at High Flyer, a hedge fund based in China, open-sourced their work on a chain-of-thought reasoning model, based on Qwen and LLama (open source LLMs).

It would be somewhat bizarre to describe Meta's open sourcing of LLama as "the Americans gifting a model", despite Meta having a corporate headquarters in the United States.

Re: Nvidia’s $589B DeepSeek rout

#409
post #383

Earlier quoted context omitted.

This is a cookie cutter comment that appears to have been copy pasted from a thread about Gamestop or something. DeepSeek R1 allegedly being almost 50x more compute efficient isn't just a "vague rumor". You do this community a disservice by commenting before understanding what investors are thinking at the current moment.

Has anyone verified DeepSeek's claims about R1? They have literally published one single paper and it has been out for a week. Nothing about what they did changed Nvidia's fundamentals. In fact there was no additional news over the weekend or today morning. The entire market movement is because of a single statement by DeepSeek's CEO from over a week ago. People sold because other people sold. This is exactly how a p…

Not a true verification but I have tried the Deepseek R1 7b model running locally, it runs on my 6gb laptop GPU and the results are impressive.

Its obviously constrained by this hardware and this model size as it does some strange things sometimes and it is slow (30 secs to respond) but I've got it to do some impressive things that GPT4 struggles with or fails on.

Also of note I asked it about Taiwan and it parroted the official CCP line about Taiwan being part of China, without even the usual delay while it generated the result.

Re: Nvidia’s $589B DeepSeek rout

#410

Earlier quoted context omitted.

Not quite, I believe this sell off was caused by DeepSeek showing with their new model that the hardware demands of AI are not necessarily as high as everyone has assumed (as required by competing models). I've tried their 7b model, running locally on a 6gb laptop GPU. Its not fast, but the results I've had have rivaled GPT4. Its impressive.

I believe you that it had to do with the selloff, but I believe that efficiency improvements are good news for NVIDIA: each card just got 20x more useful

each card is not 20x more useful lol. there's no evidence yet that the deepseek architecture would even yield a substantially (20x) more performant model with more compute.

if there's evidence to the contrary I'd love to see. in any case I don't think a h800 is even 20x better than a h100 anyway, so the 20x increase has to be wrong.

Post reply on HN