Earlier quoted context omitted.
You could even say that overconfidence in abilities and knowledge outside of core expertise is the critical flaw of today's entire tech industry.
With the crowdstrike outage earlier last year it was incredible how many hidden security and kernel "experts" came out crawling from the woodwork, questioning why anything needs to run in the kernel and predicting the company's demise.
Nvidia’s $589B DeepSeek rout
771–780 of 1001 posts
Re: Nvidia’s $589B DeepSeek rout
#772I find it interesting because the DeepSeek stuff, while very cool, doesn't seem invalidate that more compute wouldn't translate to even _higher_ capabilities? It's amazing what they did with a limited budget, but instead of the takeaway being "we don't need that much compute to achieve X", it could also be, "These new results show that we can achieve even 1000*X with our currently planned compute buildout" But perhap…
Probably not. If the price of Nvidia is dropping, it's because investors see a world where Nvidia hardware is less valuable, probably because it will be used less. You can't do the distill/magnify cycle like you do with alphago. LLM models have basically stalled in their base capabilities, pre training is basically over at this point, so the news arms race will be over marginal capability gains and (mostly) making th…
are you sure? people are saying that there’s an analogous cycle where you use o1-style reasoning to produce better inputs to the next training round
Re: Nvidia’s $589B DeepSeek rout
#773Earlier quoted context omitted.
Well you have to keep in mind that Nvidia has a 3 trillion dollar valuation. That kind of heavy valuation comes with heavy expectations about future growth. Some of those assumptions about future Nvidia growth are their ability to maintain their heavy growth rates, for very far into the future. Training is a huge component of Nvidia's projected growth. Inference is actually much more competitive, but training is almo…
My layman view is that more compute (more reasoning) will not solve harder problems. I'm using those models every day and when problem hits a certain complexity it will fail, no matter how much it "reasons"
Re: Nvidia’s $589B DeepSeek rout
#774NVIDIA sells shovels to the gold rush. One miner (Liang Wenfeng), who has previously purchased at least 10,000 A100 shovels... has a "side project" where they figured out how to dig really well with a shovel and shared their secrets. The gold rush, wether real or a bubble is still there! NVIDA will still sell every shovel they can manufacture, as soon as it is available in inventory. Fortune 100 companies will still…
Can anyone comment on why Wenfeng shared his secret sauce? Other than publicity, there only seems to be downsides for him, as now everyone else with larger compute will just copy and improve?
Re: Nvidia’s $589B DeepSeek rout
#775Here’s a take I haven’t seen yet: If training and inference just got 40x more efficient, but OpenAI and co. still have the same compute resources, once they’ve baked in all the DeepSeek improvements, we’re about to find out very quickly whether 40x the compute delivers 40x the performance / output quality, or if output quality has ceased to be compute-bound.
I've missed the stories on this until now. Is it known (and is there an ELI5) how they were able to do it so much more efficiently?
Re: Nvidia’s $589B DeepSeek rout
#776Earlier quoted context omitted.
> the Chinese Daya Guo, Dejian Yang, Haowei Zhang, et.al., quant researchers at High Flyer, a hedge fund based in China, open-sourced their work on a chain-of-thought reasoning model, based on Qwen and LLama (open source LLMs). It would be somewhat bizarre to describe Meta's open sourcing of LLama as "the Americans gifting a model", despite Meta having a corporate headquarters in the United States.
Thank you. The amount of casual sinophobia allowed on hackernews has been a real turn off. I find myself avoiding threads like these in anticipation of these comments
Re: Nvidia’s $589B DeepSeek rout
#777Earlier quoted context omitted.
I think you’re wrong and Wallstreet got Deepseek’s impact wrong. You say DeepSeek should decrease Nvidia demand. Wallstreet agreed today. I say DeepSeek should increase Nvidia’s demand due to Jevon’s Paradox.
> DeepSeek should increase Nvidia’s demand due to Jevon’s Paradox. How exactly? From what I’ve read the full model can run on MacBook M1 sort of hardware just fine. And this is their first release, I’d expect it to get more efficient and maybe domain specific models can be run on much lower grade hardware sort of raspberry pi sort.
Re: Nvidia’s $589B DeepSeek rout
#778Earlier quoted context omitted.
IIUC they released a paper, it's partially algorithmic improvements partially good old low level optimization.
There's something to be said about the idea that instead of just dumping oceans of money into buying Nvidia cards they just...optimized what they had I'd say the wider industry could learn a thing or two, but as other commentors have joked. The line must go up
Re: Nvidia’s $589B DeepSeek rout
#779Earlier quoted context omitted.
But every successful SV founder and or VC is not only a tech genius but also a geopolitical and socioeconomic expert! That’s why they make war companies, cozy up to politicians, and talk about how woke is ruining the world. /s
In fairness, 'geopolitical experts' may not really exist. There are a range of people who make up interesting stories to a greater or lesser extent but all seem to be serially misinformed. Some things are too complicated to have expertise in. Indeed, while the existence of socioeconomic experts seems more likely we don't have any way of reliably identifying them. The people who actually end up making social or econom…
Except for, I don't know, the many thousands of people who work at various government agencies (diplomatic, intelligence) or even private sector policy circles whose job it is to literally be geopolitical experts in a given area.
Re: Nvidia’s $589B DeepSeek rout
#780Earlier quoted context omitted.
> the Chinese Daya Guo, Dejian Yang, Haowei Zhang, et.al., quant researchers at High Flyer, a hedge fund based in China, open-sourced their work on a chain-of-thought reasoning model, based on Qwen and LLama (open source LLMs). It would be somewhat bizarre to describe Meta's open sourcing of LLama as "the Americans gifting a model", despite Meta having a corporate headquarters in the United States.
Thank you. The amount of casual sinophobia allowed on hackernews has been a real turn off. I find myself avoiding threads like these in anticipation of these comments