Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

51–60 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#52

Earlier quoted context omitted.

CEO of Scale said Deepseek is lying and actually has a 50k GPU cluster. He said they lied in the paper because technically they aren't supposed to have them due to export laws. I feel like this is very likely. They obvious did some great breakthroughs, but I doubt they were able to train on so much less hardware.

CEO of a human based data labelling services company feels threatened by a rival company that claims to have trained a frontier class model with an almost entirely RL based approach, with a small cold start dataset (a few thousand samples). It's in the paper. If their approach is replicated by other labs, Scale AI's business will drastically shrink or even disappear. Under such dire circumstances, lying isn't entirel…

[deleted]

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#54

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

It’s not just the economy that is vulnerable, but global geopolitics. It’s definitely worrying to see this type of technology in the hands of an authoritarian dictatorship, especially considering the evidence of censorship. See this article for a collected set of prompts and responses from DeepSeek highlighting the propaganda: https://medium.com/the-generator/deepseek-hidden-china-polit... But also the claimed cost i…

I am certainly reliefed there is no super power lock in for this stuff.

In theory I could run this one at home too without giving my data or money to Sam Altman.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#55

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

Why do americans think china is like a hivemind controlled by an omnisicient Xi, making strategic moves to undermine them? Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x?

I mean what’s also incredible about all this cope is that it’s exactly the same David-v-Goliath story that’s been lionized in the tech scene for decades now about how the truly hungry and brilliant can form startups to take out incumbents and ride their way to billions. So, if that’s not true for DeepSeek, I guess all the people who did that in the U.S. were also secretly state-sponsored operations to like make better SAAS platforms or something?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#56

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

Would you say they were more vulnerable if the PRC kept it secret so as not to disclose their edge in AI while continuing to build on it?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#57

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

CEO of Scale said Deepseek is lying and actually has a 50k GPU cluster. He said they lied in the paper because technically they aren't supposed to have them due to export laws. I feel like this is very likely. They obvious did some great breakthroughs, but I doubt they were able to train on so much less hardware.

Just to check my math: They claim something like 2.7 million H800 hours which would be less than 4000 GPU units for one month. In money something around 100 million USD give or take a few tens of millions.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#58

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

We will know soon enough if this replicates since Huggingface is working on replicating it.

To know that this would work requires insanely deep technical knowledge about state of the art computing, and the top leadership of the PRC does not have that.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#59

Genuinly curious, what is everyone using reasoning models for? (R1/o1/o3)

Regular coding questions mostly. For me o1 generally gives better code and understands the prompt more completely (haven’t started using r1 or o3 regularly enough to opine).

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#60

Even if you think this particular team cheated, the idea that nobody will find ways of making training more efficient seems silly - these huge datacenter investments for purely AI will IMHO seem very short sighted in 10 years

More like three years. Even in the best case the retained value curve of GPUs is absolutely terrible. Most of these huge investments in GPUs are going to be massive losses.
Post reply on HN