Live data from Hacker News

Nvidia’s $589B DeepSeek rout

finance.yahoo.com

291–300 of 1001 posts

Re: Nvidia’s $589B DeepSeek rout

#291
Massive overreaction in my opinion (not financial advice).

People are forgetting how much compute is being used for inference. This is going to be further accelerated:

- Reasoning models are going to generate orders of magnitude more tokens.

- Synthetic an augmented data is more and more prevalent, and if you need to process pre-training scale dataset, you will need a lot of inference.

- etc.

Re: Nvidia’s $589B DeepSeek rout

#292
post #85

Earlier quoted context omitted.

The issue here is not that DeepSeek exists as a competitor to GPT, Claude, Gemini,... The issue is that DeepSeek have shown that you don't need that much raw computing power to run an AI, which means that companies including OpenAI may focus more on efficiency than on throwing more GPUs at the problem, which is not good news for those in the business of making GPUs. At least according to the market.

One of the questions about this is that of the US’s human capital, i.e. does the US (still) have enough capable tech people in order to make that happen?

Lol, yes. The US is still very much at the forefront of this stuff. DeepSeek have presented some neat optimizations, but there have been many such papers and optimizations get implemented quickly once someone has proven them out.

Re: Nvidia’s $589B DeepSeek rout

#293
post #219

Earlier quoted context omitted.

It's a very strange result. I believe that NVIDIA is overvalued, but if DeepSeek really is as great as has been said, then it'll be even greater when scaled up to OpenAI sizes, and when you get more out you have more reason to pay, so this should if it pans out lead to more demand for GPUs-- basically Jevon's paradox.

You need to be prepared for the reality that naive scaling no longer works for LLMs anymore. Simple question: where is GPT-5?

There is a theory that Deepseek gives based on it's distillation process that hints towards, that o1 is really a distillation of a bigger GPT (GPT5?).

Some consider this to be spurious/conspiracy.

Re: Nvidia’s $589B DeepSeek rout

#294

Earlier quoted context omitted.

PRC just announced mass producing 28nm litho that cost 1/30 ASML hardware. Easy to extrapolate where this goes especially mature nodes like 28nm still accounts fo over 70% of global wafer use.

I'm increasingly believing that the West has turned their dream of free trade for comparative advantage into a massive deindustrialization. The end result is unfolding in front of everyone and the sentiment I see, even on HN, is we can't outcompete China any more. This is sad. Really sad. And this fits exactly what Liu Cixin said in Three Body Problem: Weakness and ignorance are not barriers to survival, but arroganc…

This is the kind of overly dramatic thinking that leads to stock market plunges.

China is an enormous country. It has over 4x the population of the USA. Unless you assume Chinese people are fundamentally different, it should be producing 4x the output in every field vs America. Yet the impact and legacy of communism is dire: China clearly isn't even close to 4x the productivity of the USA. How many companies on the leading edge of AI does the USA have? Meta, OpenAI, Anthropic, Google, NVIDIA, Cerebras, X.ai to pick just a handful of thousands.

Meanwhile Europe has produced one, Mistral (or two if you count DeepMind), and China has produced one. DeepSeek meanwhile, despite being impressive, has been doing the usual thing Chinese firms focus on of rapidly driving down the cost of tech already proven out by companies elsewhere. They have a long history of doing this and it's something they take cultural pride in, but at the same time, Chinese tech executives do worry about their relative lack of leading edge innovation. The head of DeepSeek has given interviews where he talks about that specifically and their desire to change attitudes and ideas about what Chinese firms can do, because there's a widespread cultural belief there that the Americans go from 0-1 and the Chinese can go from 1-10.

It's also worth remembering that prices in China are artificial. It's a somewhat planned economy still. Sectors of the economy with military relevance are heavily subsidized and they play games with their exchange rates, indeed perhaps in an attempt to forcibly deindustrialize the west. Just because something is made cheaper there doesn't necessarily mean they're doing it better. It can also be that they're just subsidized all the way to do that, and the average Chinese citizen is the loser (because they can't afford to buy things that would otherwise be affordable to them).

Re: Nvidia’s $589B DeepSeek rout

#295

Earlier quoted context omitted.

This would be an excellent explanation if DeepSeek had announced its model over the last weekend rather than weeks ago, and if R1 wasn't a COT reasoning model which needs a lot more inference time compute than other SOTA models like llama.

Information lag, especially with respect to PRC developments and technical developments. Taking 1-2 week for info to be shared and passed down info chain not unusual. IRC COT typically increase inference 1-5x depending on task complexity, i.e. instead of scaling down compute demand by 50x, it's 10x, which is still substantial. Could investors be panicking? Sure, but there's rational basis for doing so.

DeepSeek V3 is a 671B parameter MOE model? I am not sure why it's 50x cheaper at inference time than other models. We don't know what the cost of running o1 is, but I doubt it has 50x as many params as R1. Most the advantages of MOE are reduced when using reasonable branch sizes so that wouldn't make R1 cheaper in practice either. I think people might be seeing a lower markup from DeepSeek and confusing it with cheaper inference?

Re: Nvidia’s $589B DeepSeek rout

#296

The part of this that doesn’t jibe with me is the fact that they also released this incredibly detailed technical report on their architecture and training strategy. The paper is well-written and has a lot of specifics. Exactly the opposite of what you would do if you had truly made an advancement of world-altering magnitude. All this says to me is that the models themselves have very little intrinsic value / are hig…

> The part of this that doesn’t jibe with me is the fact that they also released this incredibly detailed technical report on their architecture and training strategy. The paper is well-written and has a lot of specifics. Exactly the opposite of what you would do if you had truly made an advancement of world-altering magnitude. I disagree completely on this sentiment. This was in fact the trend for a century or more…

I’m not arguing for/against the altruistic ideal of sharing technological advancements with society, I’m just saying that having a great model architecture is really not a defensible value proposition for a business. Maybe more accurate to say publishing everything in detail indicates that it’s likely not a defensible advancement, not that it isn’t significant.

Re: Nvidia’s $589B DeepSeek rout

#297

ASML plunge indicates a hysterical/irrational component to the response, right? They aren’t going anywhere. If it turns out training is easier than expected, they make the devices that make the devices that do inference too… If the field is going to produce anything useful, cheap training gets us there faster.

Not making the news in western media yet is PRC claims to have started mass producing their indigenous 28nm litho (70% of global wafer use) this month... the estimated cost is 1/30th of ASML machines. Extrapolate and they're on trend of produce 14nm machines at comparable fraction cost in next few years.

It seems hard to extrapolate there—the 30nm-10nm range is where Intel really started to start having trouble, right?

Anyway, this seems like a bigger problem for companies whose business model is actually selling those chips. It couldn’t be the case that much of ASML’s valuation is based on people continuing to use their old 28nm machines, right?

Re: Nvidia’s $589B DeepSeek rout

#298

The people writing market commentary are simply making it up. The news about DeepSeek is not new and doesn't reduce the value of ASML. People are selling now because they are scared because the number went down.

If you look at Intel, it is a cyclical business.

One could argue by extension that ASML is also cyclical.

Re: Nvidia’s $589B DeepSeek rout

#299

Earlier quoted context omitted.

Ugh. No, we can now run a decent model through CPU. Not an expensive video card. Just try out the standard deepseek-r1 or even the deepseek-r1:1.5B through ollama. No need for expensive hardware anymore locally. My PC ( without Nvidia card/expensive hardware) runs the a deepseek 1.5 b query fast enough - 2 - 9 seconds until it's finished.

I would shoot myself if I had to wait 9 seconds for a query. I sometimes even would like to kill by browser taking seconds to open a page...

2-4 seconds without streaming on a non GPU machine

Re: Nvidia’s $589B DeepSeek rout

#300

Earlier quoted context omitted.

Right and LLMs will not be able to generate their own high quality training data. There are no perpetual motion machines.

> LLMs will not be able to generate their own high quality training data. Humans certainly did. We did not inherit our physics and poetry books from some aliens.

I can't prove that we did but I don't know that we /didn't/.
Post reply on HN