Live data from Hacker News

Nvidia’s $589B DeepSeek rout

finance.yahoo.com

171–180 of 1001 posts

Re: Nvidia’s $589B DeepSeek rout

#171
post #150

Traders are saying not doing multitoken prediction, not using Sharpe ratio adjusted rewards, using reward models, and not compressing KV cache tokens by >90%, were supposed to be worth hundreds of billions of dollars of future expected revenue flow, at least according to other traders. I say to the traders: you should have just stuck to reading arxiv, TPOT, and jhana twitter for the past 2 years, rather than listenin…

> you should have just stuck to reading arxiv, TPOT, and jhana twitter for the past 2 years

Traders expect to be able to trade based on regurgitated headlines made by tech influencers who don't know ANYTHING about NOTHING.

Re: Nvidia’s $589B DeepSeek rout

#172
What I’d like to know is.. If a good model can be trained with much fewer GPUs using a breakthrough technique, can the breakthrough technique be used by OpenAI, MSFT et al who has loads of GPUs to train a model that is orders of magnitude better than their state of the art today?

We’ve been getting the impression that the limiting factor was the number of GPUs right? If so, this reduces that bottleneck and frees them up to do even better right?

Re: Nvidia’s $589B DeepSeek rout

#173

Earlier quoted context omitted.

Meta should actually go up from this. If deep seek is perfect they don't need to pay for an expensive Llama team. And even if deep seek isn't perfect, the low cost training strategies they've invented could be used by Meta to reduce the cost of Llama training. Although Meta develops models they don't sell them. So a world where foundation models are free is fine for them.

But if they don't sell them, are they spending 50 billion a year on open source goodwill?

But now (presumably) they don't need to spend 50 billion? e.g. 5 billion or whatever might be enough which makes it even easier for them to justify this.

Re: Nvidia’s $589B DeepSeek rout

#174

The part of this that doesn’t jibe with me is the fact that they also released this incredibly detailed technical report on their architecture and training strategy. The paper is well-written and has a lot of specifics. Exactly the opposite of what you would do if you had truly made an advancement of world-altering magnitude. All this says to me is that the models themselves have very little intrinsic value / are hig…

What has changed is the perception that people like OpenAI/MSFT would have an edge on the competition because of their huge datacenters full of NVDA hardware. That is no longer true. People now believe that you can build very capable AI applications for far less money. So the perception is that the big guys no longer have an edge. Tesla had already proven that to be wrong. Tesla's Hardware 3 is a 6 year old design, a…

I mean, I think they still do have an edge - ChatGPT is a great app and has strong consumer recognition already, very hard to displace.. and MSFT has a major installed base of enterprise customers who cannot readily switch cloud / productivity suite providers. So I guess they still have an edge it’s just nore of a traditional edge.

Re: Nvidia’s $589B DeepSeek rout

#175
post #2

Nvidia -13% in Frankfurt stock market just now. Valuations of private unicorns like OpenAi and Anthropic must be in free fall. DeepSeek spends $6 million in old H800 hardware to develop open source model that overtakes ChatGPT. AI gets better, but profit margins sink with strong competition. Chinese AI startup DeepSeek overtakes ChatGPT on Apple App Store https://news.ycombinator.com/item?id=42839656 Edit: Nvidia now…

> DeepSeek spends $6 million in old H800 hardware to develop open source model that overtakes ChatGPT. DeepSeek claims that's what they spent. They're under a trade embargo, and if they had access to any more than that it would have been obtained illegally. They might be telling the truth, but let's wait until someone else replicates it before we fully accept it.

All of the western AI companies trained on illegally obtained data, they barely even bother to deny it. This is an industry where lies are normalised. (Not to contradict your point about this specific number)

Re: Nvidia’s $589B DeepSeek rout

#176
post #128

Earlier quoted context omitted.

> DeepSeek spends $6 million in old H800 hardware to develop open source model that overtakes ChatGPT. DeepSeek claims that's what they spent. They're under a trade embargo, and if they had access to any more than that it would have been obtained illegally. They might be telling the truth, but let's wait until someone else replicates it before we fully accept it.

Huggingface is currently replicating it. Replications of small models indicate that they don't lie any significant amount. The architecture is cheap to train. Berkeley Researchers Replicate DeepSeek R1's Core Tech for Just $30: A Small Model RL Revolution https://xyzlabs.substack.com/p/berkeley-researchers-replicat...

Whoah, that's incredible!

I remember a year ago I was hoping that in a decade from now it would be great to run GPT4-class models on my own hardware. The reality seems to be far more exciting.

Re: Nvidia’s $589B DeepSeek rout

#177
post #2

Nvidia -13% in Frankfurt stock market just now. Valuations of private unicorns like OpenAi and Anthropic must be in free fall. DeepSeek spends $6 million in old H800 hardware to develop open source model that overtakes ChatGPT. AI gets better, but profit margins sink with strong competition. Chinese AI startup DeepSeek overtakes ChatGPT on Apple App Store https://news.ycombinator.com/item?id=42839656 Edit: Nvidia now…

It's a very strange result. I believe that NVIDIA is overvalued, but if DeepSeek really is as great as has been said, then it'll be even greater when scaled up to OpenAI sizes, and when you get more out you have more reason to pay, so this should if it pans out lead to more demand for GPUs-- basically Jevon's paradox.

The architecture doesn't keep yielding better results, Jevon's paradox doesn't apply.

Re: Nvidia’s $589B DeepSeek rout

#178
post #86
post #59

Earlier quoted context omitted.

Yours is the correct take, it isn't about who did it, it was that it was done at all. This shows how effective open source and open development can be. Best way to get a correct answer is ... https://www.phind.com/search?cache=ws3qq1xl8hj4yx1izd5oeda0

Not sure I understand how exactly open source plays a key role here in terms of project development. Looking at https://github.com/deepseek-ai , those repos have a bunch of of contributors but unless I'm wrong I don't see any significant contributions. What am I missing?

One could argue that by Meta and other companies releasing open weights and detailed got us to where we are now with R1. Even if it wasn't your race car that crossed the line first, everyone can now get a copy.

Re: Nvidia’s $589B DeepSeek rout

#179

The people writing market commentary are simply making it up. The news about DeepSeek is not new and doesn't reduce the value of ASML. People are selling now because they are scared because the number went down.

ASML paper value is determined by equipment sales from projected compute supply/demand. CHIPS building redundant global fabs = glut from more excess capacity = less future sales. Stargate = excess demand from everyone spending 100s of billions of compute = need even more fabs = more future sales. Then DeepSeek = suddenly no need for that much future compute... if number of future compute demand relative to short term fab overcapacity is going down, then it's reasonable to sell. Relatively predictable Semiconductor market cycles due to cost of capex and time to build fabs / increase new wafers output to match future demand is a thing.

Re: Nvidia’s $589B DeepSeek rout

#180

Earlier quoted context omitted.

It's a very strange result. I believe that NVIDIA is overvalued, but if DeepSeek really is as great as has been said, then it'll be even greater when scaled up to OpenAI sizes, and when you get more out you have more reason to pay, so this should if it pans out lead to more demand for GPUs-- basically Jevon's paradox.

The architecture doesn't keep yielding better results, Jevon's paradox doesn't apply.

But surely it can be scaled up, or is this compression thing something making the approach good only for small models (I haven't read the Deepseek papers (can't allocate time to it))?
Post reply on HN