Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

31–40 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#31

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

CEO of Scale said Deepseek is lying and actually has a 50k GPU cluster. He said they lied in the paper because technically they aren't supposed to have them due to export laws. I feel like this is very likely. They obvious did some great breakthroughs, but I doubt they were able to train on so much less hardware.

Alexandr Wang did not even say they lied in the paper.

Here's the interview: https://www.youtube.com/watch?v=x9Ekl9Izd38. "My understanding is that is that Deepseek has about 50000 a100s, which they can't talk about obviously, because it is against the export controls that the United States has put in place. And I think it is true that, you know, I think they have more chips than other people expect..."

Plus, how exactly did Deepseek lie. The model size, data size are all known. Calculating the number of FLOPS is an exercise in arithmetics, which is perhaps the secret Deepseek has because it seemingly eludes people.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#32

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

CEO of Scale said Deepseek is lying and actually has a 50k GPU cluster. He said they lied in the paper because technically they aren't supposed to have them due to export laws. I feel like this is very likely. They obvious did some great breakthroughs, but I doubt they were able to train on so much less hardware.

I've also read that Deepseek has released the research paper and that anyone can replicate what they did.

I feel like if that were true, it would mean they're not lying.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#34

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

Why do americans think china is like a hivemind controlled by an omnisicient Xi, making strategic moves to undermine them? Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x?

think about how big the prize is, how many people are working on it and how much has been invested (and targeted to be invested, see stargate).

And they somehow yolo it for next to nothing?

yes, it seems unlikely they did it exactly they way they're claiming they did. At the very least, they likely spent more than they claim or used existing AI API's in way that's against the terms.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#36

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

CEO of Scale said Deepseek is lying and actually has a 50k GPU cluster. He said they lied in the paper because technically they aren't supposed to have them due to export laws. I feel like this is very likely. They obvious did some great breakthroughs, but I doubt they were able to train on so much less hardware.

CEO of a human based data labelling services company feels threatened by a rival company that claims to have trained a frontier class model with an almost entirely RL based approach, with a small cold start dataset (a few thousand samples). It's in the paper. If their approach is replicated by other labs, Scale AI's business will drastically shrink or even disappear.

Under such dire circumstances, lying isn't entirely out of character for a corporate CEO.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#37

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

I don't believe that the model was trained on so few GPUs, personally, but it also doesn't matter IMO. I don't think SOTA models are moats, they seem to be more like guiding lights that others can quickly follow. The volume of research on different approaches says we're still in the early days, and it is highly likely we continue to get surprises with models and systems that make sudden, giant leaps.

Many "haters" seem to be predicting that there will be model collapse as we run out of data that isn't "slop," but I think they've got it backwards. We're in the flywheel phase now, each SOTA model makes future models better, and others catch up faster.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#38

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

It’s not just the economy that is vulnerable, but global geopolitics. It’s definitely worrying to see this type of technology in the hands of an authoritarian dictatorship, especially considering the evidence of censorship. See this article for a collected set of prompts and responses from DeepSeek highlighting the propaganda:

https://medium.com/the-generator/deepseek-hidden-china-polit...

But also the claimed cost is suspicious. I know people have seen DeepSeek claim in some responses that it is one of the OpenAI models, so I wonder if they somehow trained using the outputs of other models, if that’s even possible (is there such a technique?). Maybe that’s how the claimed cost is so low that it doesn’t make mathematical sense?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#39

Even if you think this particular team cheated, the idea that nobody will find ways of making training more efficient seems silly - these huge datacenter investments for purely AI will IMHO seem very short sighted in 10 years

Isn't it possible with more efficiency, we still want them for advanced AI capabilities we could unlock in the future?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#40
Question about the rule-based rewards (correctness and format) mentioned in the paper: Does the raw base model just expected “stumble upon“ a correct answer /correct format to get a reward and start the learning process? Are there any more details about the reward modelling?
Post reply on HN