the 32b distillation just became the default model for my home server.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
131–140 of 1001 posts
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#132The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.
If someone gets something to work with 1k h100s that should have taken 100k h100s, that means the group with the 100k is about to have a much, much better model.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#133Reddit's /r/chatgpt subreddit is currently heavily brigaded by bots/shills praising r1, I'd be very suspicious of any claims about it.
I do believe they were honest in the paper, but the $5.5m training cost (for v3) is defined in a limited way: only the GPU cost at $2/hr for the one training run they did that resulted in the final V3 model. Headcount, overhead, experimentation, and R&D trial costs are not included. The paper had something like 150 people on it, so obviously total costs are quite a bit higher than the limited scope cost they disclosed, and also they didn't disclose R1 costs.
Still, though, the model is quite good, there are quite a few independent benchmarks showing it's pretty competent, and it definitely passes the smell test in actual use (unlike many of Microsoft's models which seem to be gamed on benchmarks).
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#134The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.
Correct me if I'm wrong, but couldn't you take the optimization and tricks for training, inference, etc. from this model and apply to the Big Corps' huge AI data centers and get an even better model?
I'll preface this by saying, better and better models may not actually unlock the economic value they are hoping for. It might be a thing where the last 10% takes 90% of the effort so to speak
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#135The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.
CEO of Scale said Deepseek is lying and actually has a 50k GPU cluster. He said they lied in the paper because technically they aren't supposed to have them due to export laws. I feel like this is very likely. They obvious did some great breakthroughs, but I doubt they were able to train on so much less hardware.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#136Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#137Very small training set!
"we replicate the DeepSeek-R1-Zero and DeepSeek-R1 training on small models with limited data. We show that long Chain-of-Thought (CoT) and self-reflection can emerge on a 7B model with only 8K MATH examples, and we achieve surprisingly strong results on complex mathematical reasoning. Importantly, we fully open-source our training code and details to the community to inspire more works on reasoning."
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#138Even if you think this particular team cheated, the idea that nobody will find ways of making training more efficient seems silly - these huge datacenter investments for purely AI will IMHO seem very short sighted in 10 years
More like three years. Even in the best case the retained value curve of GPUs is absolutely terrible. Most of these huge investments in GPUs are going to be massive losses.
source?
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#139Earlier quoted context omitted.
Alexandr Wang did not even say they lied in the paper. Here's the interview: https://www.youtube.com/watch?v=x9Ekl9Izd38 . "My understanding is that is that Deepseek has about 50000 a100s, which they can't talk about obviously, because it is against the export controls that the United States has put in place. And I think it is true that, you know, I think they have more chips than other people expect..." Plus, how ex…
> Plus, how exactly did Deepseek lie. The model size, data size are all known. Calculating the number of FLOPS is an exercise in arithmetics, which is perhaps the secret Deepseek has because it seemingly eludes people. Model parameter count and training set token count are fixed. But other things such as epochs are not. In the same amount of time, you could have 1 epoch or 100 epochs depending on how many GPUs you ha…
I don't expect a #180 AUM hedgefund to have as many GPUs than meta, msft or Google.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#140The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.
How likely is this? Just a cursory probing of deepseek yields all kinds of censoring of topics. Isn't it just as likely Chinese sponsors of this have incentivized and sponsored an undercutting of prices so that a more favorable LLM is preferred on the market? Think about it, this is something they are willing to do with other industries. And, if LLMs are going to be engineering accelerators as the world believes, the…