Live data from Hacker News

Nvidia’s $589B DeepSeek rout

finance.yahoo.com

311–320 of 1001 posts

Re: Nvidia’s $589B DeepSeek rout

#311

The part of this that doesn’t jibe with me is the fact that they also released this incredibly detailed technical report on their architecture and training strategy. The paper is well-written and has a lot of specifics. Exactly the opposite of what you would do if you had truly made an advancement of world-altering magnitude. All this says to me is that the models themselves have very little intrinsic value / are hig…

Build trust by releasing your inferior product for free and as open as possible. Get attention, then release your superior product behind paywall. Name recognition is incredibly important within and outside of China.

Keep in mind, they’re still competing with Baidu, Tencent and other AI labs.

Re: Nvidia’s $589B DeepSeek rout

#312
post #291

Massive overreaction in my opinion (not financial advice). People are forgetting how much compute is being used for inference. This is going to be further accelerated: - Reasoning models are going to generate orders of magnitude more tokens. - Synthetic an augmented data is more and more prevalent, and if you need to process pre-training scale dataset, you will need a lot of inference. - etc.

No one will ever need more than 640k tokens.

Re: Nvidia’s $589B DeepSeek rout

#313
post #150

Traders are saying not doing multitoken prediction, not using Sharpe ratio adjusted rewards, using reward models, and not compressing KV cache tokens by >90%, were supposed to be worth hundreds of billions of dollars of future expected revenue flow, at least according to other traders. I say to the traders: you should have just stuck to reading arxiv, TPOT, and jhana twitter for the past 2 years, rather than listenin…

Who is jhana?

It’s a part of twitter that is nominally about advanced forms of meditation but is also a surprisingly overrepresented by ai researchers lol

Re: Nvidia’s $589B DeepSeek rout

#314

The people writing market commentary are simply making it up. The news about DeepSeek is not new and doesn't reduce the value of ASML. People are selling now because they are scared because the number went down.

>The people writing market commentary are simply making it up.

Everyone except for you, right?

Re: Nvidia’s $589B DeepSeek rout

#315

The part of this that doesn’t jibe with me is the fact that they also released this incredibly detailed technical report on their architecture and training strategy. The paper is well-written and has a lot of specifics. Exactly the opposite of what you would do if you had truly made an advancement of world-altering magnitude. All this says to me is that the models themselves have very little intrinsic value / are hig…

I always thought AMZN is the winner since I looked into Bedrock. When I saw Claude on there it added a fuck yeah, and now the best models being open just takes it to another level.

AMZN: no horse picked, we host anything

MSFT: Open AI

GOOGLE: Google AI

AMZN is in the strongest position.

Re: Nvidia’s $589B DeepSeek rout

#316

The part of this that doesn’t jibe with me is the fact that they also released this incredibly detailed technical report on their architecture and training strategy. The paper is well-written and has a lot of specifics. Exactly the opposite of what you would do if you had truly made an advancement of world-altering magnitude. All this says to me is that the models themselves have very little intrinsic value / are hig…

What has changed is the perception that people like OpenAI/MSFT would have an edge on the competition because of their huge datacenters full of NVDA hardware. That is no longer true. People now believe that you can build very capable AI applications for far less money. So the perception is that the big guys no longer have an edge. Tesla had already proven that to be wrong. Tesla's Hardware 3 is a 6 year old design, a…

The perception only makes sense if it is "that's it, pack up your stall" for AI.

I think what really happened is day to day trading noise. Nothing fundamentally changed, but traders believed other people believed it would.

Re: Nvidia’s $589B DeepSeek rout

#317
post #150

Traders are saying not doing multitoken prediction, not using Sharpe ratio adjusted rewards, using reward models, and not compressing KV cache tokens by >90%, were supposed to be worth hundreds of billions of dollars of future expected revenue flow, at least according to other traders. I say to the traders: you should have just stuck to reading arxiv, TPOT, and jhana twitter for the past 2 years, rather than listenin…

You need to be very solvent to stay rational.

Re: Nvidia’s $589B DeepSeek rout

#318

Earlier quoted context omitted.

One of the questions about this is that of the US’s human capital, i.e. does the US (still) have enough capable tech people in order to make that happen?

Lol, yes. The US is still very much at the forefront of this stuff. DeepSeek have presented some neat optimizations, but there have been many such papers and optimizations get implemented quickly once someone has proven them out.

> The US is still very much at the forefront of this stuff

Doesn't look like it, because the some of the biggest US tech companies now active (including Meta and Alphabet) couldn't come up with what this much-smaller Chinese company has. Which begs the question, what is that companies like Meta, Alphabet and the like do with the (already) hundreds of billions of dollars that they invested in this space?

Re: Nvidia’s $589B DeepSeek rout

#319
post #211

Earlier quoted context omitted.

o1 also does this.

o1 does not show the reasoning trace at this point. You may be confusing the final answer for the reasoning trace in the middle, it's shown pretty clearly on r1.

I wasn't really referring much to the UI as I was the fact that it does it to begin with. The thinking in deepseek trails off into its own nonsense before it answers, whereas I feel openai's is way more structured.

Re: Nvidia’s $589B DeepSeek rout

#320

The part of this that doesn’t jibe with me is the fact that they also released this incredibly detailed technical report on their architecture and training strategy. The paper is well-written and has a lot of specifics. Exactly the opposite of what you would do if you had truly made an advancement of world-altering magnitude. All this says to me is that the models themselves have very little intrinsic value / are hig…

You are not taking into account why people are willing to pay exceedingly high prices for GPUs now and that the underlying reason may have been taken away.
Post reply on HN