Live data from Hacker News

Nvidia’s $589B DeepSeek rout

finance.yahoo.com

911–920 of 1001 posts

Re: Nvidia’s $589B DeepSeek rout

#911

Nvidia doesn't have a monopoly on GEMM and GEMV. There will be dozens of hardware vendors. It is TSMC that should be the most valuable company in the world.

By that logic TSMC is not the only one capable of 5nm. ASML should be the most valuable company in the world?

Certainly one of the most valuable, but I would still say TSMC as there are lots of other steps in the production besides photolithography (etching, ion implantation, vapor deposition, packing, ...).

Re: Nvidia’s $589B DeepSeek rout

#912

Earlier quoted context omitted.

In the long run (which in the AI world is probably ~1 year) this is very good for Nvidia, very good for the hyperscalers, and very good for anyone building AI applications. The only thing it's not good for is the idea that OpenAI and/or Anthropic will eventually become profitable companies with market caps that exceed Apple's by orders of magnitude. Oh no, anyway.

Can you guys explain what this would be bad for the OpenAI and Anthropic of the world? Wasn't the story always outlined to be we build better and better models, then we eventually get to AGI, AGI works on building better and better models even faster, and we eventually get to super AGI, which can work on building better and better models even faster... Isn't "super-optimization"(in the widest sense) what we expect to…

Because building a frontier model is expensive. But building a model as good as an existing frontier model is cheap (re: distillation).

https://en.m.wikipedia.org/wiki/Knowledge_distillation

So the takeaway is they have no moat

Re: Nvidia’s $589B DeepSeek rout

#913

Earlier quoted context omitted.

This is partially why Apple is the one that stands to gain more, and it showed. Their "small models, on device" approach can only be perfected with something like DeepSeek, and they're not exposed to NVIDIA pricing, nor have to prove investors that their approach is still valid.

Until AGI removes the need for iOS. Apple is not immune to AI disruption. The Rabbit R1 was a scam but the concept was the right approach. It was just 5 years too early. You don’t need an iPhone with AGI. You just need a 5G device with a screen, connected to an AGI.

> You don’t need an iPhone with AGI. You just need a 5G device with a screen, connected to an AGI.

no the fuck i don't need that. how do you know what i need?

Re: Nvidia’s $589B DeepSeek rout

#914
post #431

NVIDIA sells shovels to the gold rush. One miner (Liang Wenfeng), who has previously purchased at least 10,000 A100 shovels... has a "side project" where they figured out how to dig really well with a shovel and shared their secrets. The gold rush, wether real or a bubble is still there! NVIDA will still sell every shovel they can manufacture, as soon as it is available in inventory. Fortune 100 companies will still…

The gold rush is over because pre-trained models don't improve as much anymore. The application layer has massive gains in cost-to-value performance. We also gain more trust from the consumer as models don't hallucinate as much. This is what DeepSeek R1 has shown us. As Ilya Sutskever said, pre-training is now over.

We now have very expensive Nvidia shovels that use a lot of power but do very little improvement to the models.

Re: Nvidia’s $589B DeepSeek rout

#915

Here’s a take I haven’t seen yet: If training and inference just got 40x more efficient, but OpenAI and co. still have the same compute resources, once they’ve baked in all the DeepSeek improvements, we’re about to find out very quickly whether 40x the compute delivers 40x the performance / output quality, or if output quality has ceased to be compute-bound.

> If training and inference just got 40x more efficient Did training and inference just get 40x more efficient, or just training? They trained a model with impressive outputs on a limited number of GPUs, but DeepSeek is still a big model that requires a lot of resources to run. Moreover, which costs more, training a model once or using it for inference across a hundred million people multiple times a day for a year?…

I don’t think that it got more efficient. It’s that smaller models can train via larger ones cheaply. Think teacher/student relationship

https://en.m.wikipedia.org/wiki/Knowledge_distillation

Re: Nvidia’s $589B DeepSeek rout

#916

Earlier quoted context omitted.

In the long run (which in the AI world is probably ~1 year) this is very good for Nvidia, very good for the hyperscalers, and very good for anyone building AI applications. The only thing it's not good for is the idea that OpenAI and/or Anthropic will eventually become profitable companies with market caps that exceed Apple's by orders of magnitude. Oh no, anyway.

Can you guys explain what this would be bad for the OpenAI and Anthropic of the world? Wasn't the story always outlined to be we build better and better models, then we eventually get to AGI, AGI works on building better and better models even faster, and we eventually get to super AGI, which can work on building better and better models even faster... Isn't "super-optimization"(in the widest sense) what we expect to…

I believe in general the business model of building frontier models has not been fully baked out yet. Lets ignore the thought of AGI and just say models do continue to improve. In OpenAIs case they have raised lots of capital in the hopes of dominating the market. That capital pegged them at a valuation. Now you have a company with ~100 employees and supposedly a lot less capital come in a get close to OpenAIs current leading model. It has the potential to pop their balloon massively.

By releasing a lot of it opensource everyone has their hands on it. Opens the door to new companies.

Or a simple mental model, there has been this ability for third parties to get quite close to leading frontier models. The leading frontier models takes hundreds of millions of dollars and if someone is able to copy it within a years time for significantly less capital, its going to be hard game of cat and mouse.

Re: Nvidia’s $589B DeepSeek rout

#917
I feel there’s a gap missing in this thread (or I may be the one missing it)

DeepSeek proved knowledge distillation works very well and cheaply https://en.m.wikipedia.org/wiki/Knowledge_distillation

But they didn’t show how to build a new frontier model cheaply.

So, you still need massive investments to build new frontier models. But the bad part, is they can be replicated cheaply

Re: Nvidia’s $589B DeepSeek rout

#918

Earlier quoted context omitted.

Can you guys explain what this would be bad for the OpenAI and Anthropic of the world? Wasn't the story always outlined to be we build better and better models, then we eventually get to AGI, AGI works on building better and better models even faster, and we eventually get to super AGI, which can work on building better and better models even faster... Isn't "super-optimization"(in the widest sense) what we expect to…

Because building a frontier model is expensive. But building a model as good as an existing frontier model is cheap (re: distillation). https://en.m.wikipedia.org/wiki/Knowledge_distillation So the takeaway is they have no moat

Beautiful and concise, much better than my word salad.

Re: Nvidia’s $589B DeepSeek rout

#919

Earlier quoted context omitted.

In the long run (which in the AI world is probably ~1 year) this is very good for Nvidia, very good for the hyperscalers, and very good for anyone building AI applications. The only thing it's not good for is the idea that OpenAI and/or Anthropic will eventually become profitable companies with market caps that exceed Apple's by orders of magnitude. Oh no, anyway.

Can you guys explain what this would be bad for the OpenAI and Anthropic of the world? Wasn't the story always outlined to be we build better and better models, then we eventually get to AGI, AGI works on building better and better models even faster, and we eventually get to super AGI, which can work on building better and better models even faster... Isn't "super-optimization"(in the widest sense) what we expect to…

First of all, we need to just stop talking about AGI and Superintelligence. It's a total distraction from the actual value that has already been created by AI/ML over the years and will continue to be created.

That said, you have to distinguish between "good for the field of AI, the AI industry overall, and users of AI" from "good for a couple of companies that want to be the sole provider of SOTA models and extract maximum value from everyone else to drive their own equity valuations to the moon". Deepseek is positive for the former and negative for the latter.

Re: Nvidia’s $589B DeepSeek rout

#920

Earlier quoted context omitted.

IIUC they released a paper, it's partially algorithmic improvements partially good old low level optimization.

There's something to be said about the idea that instead of just dumping oceans of money into buying Nvidia cards they just...optimized what they had I'd say the wider industry could learn a thing or two, but as other commentors have joked. The line must go up

Ah but why care about efficiency when you have basically unlimited investor money?

Reminds me of Japanese cars and the OPEC boycott in the 1970s...

Post reply on HN