Live data from Hacker News

Nvidia’s $589B DeepSeek rout

finance.yahoo.com

221–230 of 1001 posts

Re: Nvidia’s $589B DeepSeek rout

#221
post #2

Nvidia -13% in Frankfurt stock market just now. Valuations of private unicorns like OpenAi and Anthropic must be in free fall. DeepSeek spends $6 million in old H800 hardware to develop open source model that overtakes ChatGPT. AI gets better, but profit margins sink with strong competition. Chinese AI startup DeepSeek overtakes ChatGPT on Apple App Store https://news.ycombinator.com/item?id=42839656 Edit: Nvidia now…

Let's see how the open vs close ecosystem wins in AI

Re: Nvidia’s $589B DeepSeek rout

#223

The part of this that doesn’t jibe with me is the fact that they also released this incredibly detailed technical report on their architecture and training strategy. The paper is well-written and has a lot of specifics. Exactly the opposite of what you would do if you had truly made an advancement of world-altering magnitude. All this says to me is that the models themselves have very little intrinsic value / are hig…

> The part of this that doesn’t jibe with me is the fact that they also released this incredibly detailed technical report on their architecture and training strategy. The paper is well-written and has a lot of specifics. Exactly the opposite of what you would do if you had truly made an advancement of world-altering magnitude.

I disagree completely on this sentiment. This was in fact the trend for a century or more (see inventions ranging from the polio vaccine to "Attention is all you need" by Vaswani et. al.) before "Open"AI became the biggest player on the market due and Sam Altman tried to bag all the gains for himself. Hopefully, we can reverse course on this trend and go back to when world-changing innovations are shared openly so they can actually change the world.

Re: Nvidia’s $589B DeepSeek rout

#225
post #211
post #187

Earlier quoted context omitted.

Deepseek shows you details of the reasoning so you can trust its answers more and you can correct it when it makes a bad turn.

o1 also does this.

o1 does not show the reasoning trace at this point. You may be confusing the final answer for the reasoning trace in the middle, it's shown pretty clearly on r1.

Re: Nvidia’s $589B DeepSeek rout

#226

Seems bad! To me this highlights a failure in the US tech industry. Silicon Valley is theoretically a triumph of entrepreneurship, where the best and best thinking wins. In reality investors picked OpenAI as the winner right out of the gate, gave it more funding than most of us can imagine so that it could dominate the market, sat back and told themselves job done. Meanwhile in another tech industry a startup had to…

And this is why competition is a good thing. Silicon Valley keeps trying to create monopolies for the purpose of maximal profit extraction to the detriment of American technological development, or they use money and brute force as a substitute for actual innovation and get worse results.

Re: Nvidia’s $589B DeepSeek rout

#227

Earlier quoted context omitted.

>been obtained illegally. PRC companies breaking US export control laws is legal (for PRC companies). Maybe they're trying to avoid US entity listing, lot's of PRC companies keep mum about growing capabilites to do so. But the mere fact Deepseek is publicizing means they're unlikely to care about the political heat that is coming and the ramifications. If anything, getting on US entity list probably locks in their em…

> PRC companies breaking US export control laws is legal So long as they don't plan to do any business with the US or any of their allies I guess.

Hard to think they plan to, PRC strategic companies that gets competitive gets entity listed anyway. And CEO seems mission driven for AGI - if US going to limit hardware inevitably then nothing to do but go gloves off, and try to dunk on competition. At this point US can take deep seek off appstores but what's the point except to look petty. Eitherway, more technical ppl have pointed out some of the R1 optimizations _only_ make sense if Deepseek was constrained to older hardware, i.e. engineer at PTX level to circumvent H800 limitations to perfrom more like H100s.

Throwing this model out also gives US allies soverign AI a launchpad... reducing US dependency is step 1 to not being US allies.

Re: Nvidia’s $589B DeepSeek rout

#228
post #150

Traders are saying not doing multitoken prediction, not using Sharpe ratio adjusted rewards, using reward models, and not compressing KV cache tokens by >90%, were supposed to be worth hundreds of billions of dollars of future expected revenue flow, at least according to other traders. I say to the traders: you should have just stuck to reading arxiv, TPOT, and jhana twitter for the past 2 years, rather than listenin…

Who is jhana?

Re: Nvidia’s $589B DeepSeek rout

#229
post #219

Earlier quoted context omitted.

It's a very strange result. I believe that NVIDIA is overvalued, but if DeepSeek really is as great as has been said, then it'll be even greater when scaled up to OpenAI sizes, and when you get more out you have more reason to pay, so this should if it pans out lead to more demand for GPUs-- basically Jevon's paradox.

You need to be prepared for the reality that naive scaling no longer works for LLMs anymore. Simple question: where is GPT-5?

If we expect that the demand for GPT-5 in AI compute is 100x of that of GPT-4 then if GPT-4 was trained in months on 10k of H100 then you would need years with 100k of H100 or maybe again months with 100k of GB200.

See, there is your answer. The issue is the compute of GPUs is way to low yet for GPT-5 if they continue parameter scaling as they used to do.

GPT3 took months on 10k A100s. 10k H100 would have done it in a fraction of a time. Blackwell could train GPT4 in 10 days with same amount of GPUs as Hopper which took months.

Don't forget GPT3 is just 2.5 years old. Training is obviously waiting for the next step up in large clusters of training speed increasement. Don't be fooled, the 2x Blackwell vs. Hopper is only chip vs. chip. 10k of Blackwell including all networking speedup is easily 10x or more faster than the same amount of Hopper. So building a 1 million Blackwell cluster means 100x more training compute compared to a 100k Hopper cluster.

Nobody starts a model training if it takes years to finish... too much risk in that.

Transfer model was introduced in 2017 and ChatGPT came out 2022. Why? Because they would have needed millions of Volta GPUs instead of thousands of Ampere GPUs to train it.

Re: Nvidia’s $589B DeepSeek rout

#230
post #186

Earlier quoted context omitted.

It's a very strange result. I believe that NVIDIA is overvalued, but if DeepSeek really is as great as has been said, then it'll be even greater when scaled up to OpenAI sizes, and when you get more out you have more reason to pay, so this should if it pans out lead to more demand for GPUs-- basically Jevon's paradox.

The limit is high quality data, not compute.

Right and LLMs will not be able to generate their own high quality training data.

There are no perpetual motion machines.

Post reply on HN