Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

101–110 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#101
post #58

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

We will know soon enough if this replicates since Huggingface is working on replicating it. To know that this would work requires insanely deep technical knowledge about state of the art computing, and the top leadership of the PRC does not have that.

Researchers from TikTok claim they already replicated it

https://x.com/sivil_taram/status/1883184784492666947?t=NzFZj...

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#102

Even if you think this particular team cheated, the idea that nobody will find ways of making training more efficient seems silly - these huge datacenter investments for purely AI will IMHO seem very short sighted in 10 years

More like three years. Even in the best case the retained value curve of GPUs is absolutely terrible. Most of these huge investments in GPUs are going to be massive losses.

Do we have any idea how long a cloud provider needs to rent them out for to make back their investment? I’d be surprised if it was more than a year, but that is just a wild guess.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#103

I've always been leery about outrageous GPU investments, at some point I'll dig through and find my prior comments where I've said as much to that effect. The CEOs, upper management, and governments derive their importance on how much money they can spend - AI gave them the opportunity for them to confidently say that if you give me $X I can deliver Y and they turn around and give that money to NVidia. The problem wa…

Agree. The "need to build new buildings, new power plants, buy huge numbers of today's chips from one vendor" never made any sense considering we don't know what would be done in those buildings in 5 years when they're ready.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#104

I've always been leery about outrageous GPU investments, at some point I'll dig through and find my prior comments where I've said as much to that effect. The CEOs, upper management, and governments derive their importance on how much money they can spend - AI gave them the opportunity for them to confidently say that if you give me $X I can deliver Y and they turn around and give that money to NVidia. The problem wa…

I agree. I think there’s a good chance that politicians & CEOs pushing for 100s of billions spent on AI infrastructure are going to look foolish.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#105

Even if you think this particular team cheated, the idea that nobody will find ways of making training more efficient seems silly - these huge datacenter investments for purely AI will IMHO seem very short sighted in 10 years

Isn't it possible with more efficiency, we still want them for advanced AI capabilities we could unlock in the future?

Operating costs are usually a pretty significant factor in total costs for a data center. Unless power efficiency stops improving much and/or demand so far outstrips supply that they can't be replaced, a bunch of 10 year old GPUs probably aren't going to be worth running regardless.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#107
post #31

Earlier quoted context omitted.

CEO of Scale said Deepseek is lying and actually has a 50k GPU cluster. He said they lied in the paper because technically they aren't supposed to have them due to export laws. I feel like this is very likely. They obvious did some great breakthroughs, but I doubt they were able to train on so much less hardware.

Alexandr Wang did not even say they lied in the paper. Here's the interview: https://www.youtube.com/watch?v=x9Ekl9Izd38 . "My understanding is that is that Deepseek has about 50000 a100s, which they can't talk about obviously, because it is against the export controls that the United States has put in place. And I think it is true that, you know, I think they have more chips than other people expect..." Plus, how ex…

> Plus, how exactly did Deepseek lie. The model size, data size are all known. Calculating the number of FLOPS is an exercise in arithmetics, which is perhaps the secret Deepseek has because it seemingly eludes people.

Model parameter count and training set token count are fixed. But other things such as epochs are not.

In the same amount of time, you could have 1 epoch or 100 epochs depending on how many GPUs you have.

Also, what if their claim on GPU count is accurate, but they are using better GPUs they aren't supposed to have? For example, they claim 1,000 GPUs for 1 month total. They claim to have H800s, but what if they are using illegal H100s/H200s, B100s, etc? The GPU count could be correct, but their total compute is substantially higher.

It's clearly an incredible model, they absolutely cooked, and I love it. No complaints here. But the likelihood that there are some fudged numbers is not 0%. And I don't even blame them, they are likely forced into this by US exports laws and such.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#108
post #26

Earlier quoted context omitted.

I would think the CEO of an American AI company has every reason to neg and downplay foreign competition... And since it's a businessperson they're going to make it sound as cute and innocuous as possible

If we're going to play that card, couldn't we also use the "Chinese CEO has every reason to lie and say they did something 100x more efficient than the Americans" card? I'm not even saying they did it maliciously, but maybe just to avoid scrutiny on GPUs they aren't technically supposed to have? I'm thinking out loud, not accusing anyone of anything.

Then the question becomes, who sold the GPUs to them? They are supposedly scarse and every player in the field is trying to get ahold as many as they can, before anyone else in fact.

Something makes little sense in the accusations here.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#109
post #35

Reddit's /r/chatgpt subreddit is currently heavily brigaded by bots/shills praising r1, I'd be very suspicious of any claims about it.

The counternarrative is that it is a very accomplished piece of work that most in the sector were not expecting -- it's open source with API available at fraction of comparable service cost

It has upended a lot of theory around how much compute is likely needed over next couple of years, how much profit potential the AI model vendors have in nearterm and how big an impact export controls are having on China

V3 took top slot on HF trending models for first part of Jan ... r1 has 4 of the top 5 slots tonight

Almost every commentator is talking about nothing else

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#110
post #75

Earlier quoted context omitted.

China is actually just one person (Xi) acting in perfect unison and its purpose is not to benefit its own people, but solely to undermine the West.

If China is undermining the West by lifting up humanity, for free, while ProprietaryAI continues to use closed source AI for censorship and control, then go team China. There's something wrong with the West's ethos if we think contributing significantly to the progress of humanity is malicious. The West's sickness is our own fault; we should take responsibility for our own disease, look critically to understand its r…

> There's something wrong with the West's ethos if we think contributing significantly to the progress of humanity is malicious.

Who does this?

The criticism is aimed at the dictatorship and their politics. Not their open source projects. Both things can exist at once. It doesn't make China better in any way. Same goes for their "radical cures" as you call it. I'm sure Uyghurs in China would not give a damn about AI.

Post reply on HN