Even if you think this particular team cheated, the idea that nobody will find ways of making training more efficient seems silly - these huge datacenter investments for purely AI will IMHO seem very short sighted in 10 years
More like three years. Even in the best case the retained value curve of GPUs is absolutely terrible. Most of these huge investments in GPUs are going to be massive losses.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
91–100 of 1001 posts
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#92Earlier quoted context omitted.
CEO of Scale said Deepseek is lying and actually has a 50k GPU cluster. He said they lied in the paper because technically they aren't supposed to have them due to export laws. I feel like this is very likely. They obvious did some great breakthroughs, but I doubt they were able to train on so much less hardware.
I would think the CEO of an American AI company has every reason to neg and downplay foreign competition... And since it's a businessperson they're going to make it sound as cute and innocuous as possible
I'm not even saying they did it maliciously, but maybe just to avoid scrutiny on GPUs they aren't technically supposed to have? I'm thinking out loud, not accusing anyone of anything.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#93Earlier quoted context omitted.
Why do americans think china is like a hivemind controlled by an omnisicient Xi, making strategic moves to undermine them? Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x?
China is actually just one person (Xi) acting in perfect unison and its purpose is not to benefit its own people, but solely to undermine the West.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#94Even if you think this particular team cheated, the idea that nobody will find ways of making training more efficient seems silly - these huge datacenter investments for purely AI will IMHO seem very short sighted in 10 years
There is a big balloon full of AI hype going up right now, and regrettably it may need those data-centers. But I'm hoping that if the worst (the best) comes to happen, we will find worthy things to do with all of that depreciated compute. Drug discovery comes to mind.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#95Earlier quoted context omitted.
> Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x They literally published all their methodology. It's nothing groundbreaking, just western labs seem slow to adopt new research. Mixture of experts, key-value cache compression, multi-token prediction, 2/3 of these weren't invented by DeepSeek. They did invent a new hardware-aware distributed training approach for mixture…
"nothing groundbreaking" It's extremely cheap, efficient and kicks the ass of the leader of the market, while being under sanctions with AI hardware. Most of all, can be downloaded for free, can be uncensored, and usable offline. China is really good at tech, it has beautiful landscapes, etc. It has its own political system, but to be fair, in some way it's all our future. A bit of a dystopian future, like it was in…
So yes, DeepSeek-R1 appears to be not even be best in class, merely best open source. The only sense in which it is "leading the market" appears to be the sense in which "free stuff leads over proprietary stuff". Which is true and all, but not a groundbreaking technical achievement.
The DeepSeek-R1 distilled models on the other hand might actually be leading at something... but again hard to say it's groundbreaking when it's combining what we know we can do (small models like llama) with what we know we can do (thinking models).
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#96I've always been leery about outrageous GPU investments, at some point I'll dig through and find my prior comments where I've said as much to that effect. The CEOs, upper management, and governments derive their importance on how much money they can spend - AI gave them the opportunity for them to confidently say that if you give me $X I can deliver Y and they turn around and give that money to NVidia. The problem wa…
Latest GPUs and efficiency are not mutually exclusive, right? If you combine them both presumably you can build even more powerful models.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#97I've always been leery about outrageous GPU investments, at some point I'll dig through and find my prior comments where I've said as much to that effect. The CEOs, upper management, and governments derive their importance on how much money they can spend - AI gave them the opportunity for them to confidently say that if you give me $X I can deliver Y and they turn around and give that money to NVidia. The problem wa…
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#98The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.
From what I've read, DeepSeek is a "side project" at a Chinese quant fund. They had the GPU capacity to spare.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#99I've always been leery about outrageous GPU investments, at some point I'll dig through and find my prior comments where I've said as much to that effect. The CEOs, upper management, and governments derive their importance on how much money they can spend - AI gave them the opportunity for them to confidently say that if you give me $X I can deliver Y and they turn around and give that money to NVidia. The problem wa…
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#100Earlier quoted context omitted.
China is actually just one person (Xi) acting in perfect unison and its purpose is not to benefit its own people, but solely to undermine the West.
This explains so much. It’s just malice, then? Or some demonic force of evil? What does Occam’s razor suggest? Oh dear