Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

141–150 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#141

Earlier quoted context omitted.

CEO of Scale said Deepseek is lying and actually has a 50k GPU cluster. He said they lied in the paper because technically they aren't supposed to have them due to export laws. I feel like this is very likely. They obvious did some great breakthroughs, but I doubt they were able to train on so much less hardware.

I've also read that Deepseek has released the research paper and that anyone can replicate what they did. I feel like if that were true, it would mean they're not lying.

You can't replicate it exactly because you don't know their dataset or what exactly several of their proprietary optimizations were

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#142
post #35

Reddit's /r/chatgpt subreddit is currently heavily brigaded by bots/shills praising r1, I'd be very suspicious of any claims about it.

Its pretty nutty indeed. The model still might be good, but the botting is wild. On that note, one of my favorite benchmarks to watch is simple bench and R! doesn't perform as well on that benchmark as all the other public benchmarks, so it might be telling of something.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#143

I've always been leery about outrageous GPU investments, at some point I'll dig through and find my prior comments where I've said as much to that effect. The CEOs, upper management, and governments derive their importance on how much money they can spend - AI gave them the opportunity for them to confidently say that if you give me $X I can deliver Y and they turn around and give that money to NVidia. The problem wa…

Such a good comment.

Remember when Sam Altman was talking about raising 5 trillion dollars for hardware?

insanity, total insanity.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#144
I don’t think this entirely invalidates massive GPU spend just yet:

“ Therefore, we can draw two conclusions: First, distilling more powerful models into smaller ones yields excellent results, whereas smaller models relying on the large-scale RL mentioned in this paper require enormous computational power and may not even achieve the performance of distillation. Second, while distillation strategies are both economical and effective, advancing beyond the boundaries of intelligence may still require more powerful base models and larger-scale reinforcement learning.”

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#145
post #23

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

or maybe the US economy will do even better because more people will be able to use AI at a low cost. OpenAI will be also be able to serve o3 at a lower cost if Deepseek had some marginal breakthrough OpenAI did not already think of.

I think this is the most productive mindset. All of the costs thus far are sunk, the only move forward is to learn and adjust.

This is a net win for nearly everyone.

The world needs more tokens and we are learning that we can create higher quality tokens with fewer resources than before.

Finger pointing is a very short term strategy.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#146
Larry Ellison is 80. Masayoshi Son is 67. Both have said that anti-aging and eternal life is one of their main goals with investing toward ASI.

For them it's worth it to use their own wealth and rally the industry to invest $500 billion in GPUs if that means they will get to ASI 5 years faster and ask the ASI to give them eternal life.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#147

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

It’s not just the economy that is vulnerable, but global geopolitics. It’s definitely worrying to see this type of technology in the hands of an authoritarian dictatorship, especially considering the evidence of censorship. See this article for a collected set of prompts and responses from DeepSeek highlighting the propaganda: https://medium.com/the-generator/deepseek-hidden-china-polit... But also the claimed cost i…

> It’s definitely worrying to see this type of technology in the hands of an authoritarian dictatorship

What do you think they will do with the AI that worries you? They already had access to Llama, and they could pay for access to the closed source AIs. It really wouldn't be that hard to pay for and use what's commercially available as well, even if there is embargo or whatever, for digital goods and services that can easily be bypassed

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#148
post #131

I'm impressed by not only how good deepseek r1 is, but also how good the smaller distillations are. qwen-based 7b distillation of deepseek r1 is a great model too. the 32b distillation just became the default model for my home server.

How much VRAM is needed for the 32B distillation?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#149
post #125

"Reasoning" will be disproven for this again within a few days I guess. Context: o1 does not reason, it pattern matches. If you rename variables, suddenly it fails to solve the request.

reasoning is pattern matching at a certain level of abstraction.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#150
post #35

Reddit's /r/chatgpt subreddit is currently heavily brigaded by bots/shills praising r1, I'd be very suspicious of any claims about it.

The amount of astroturfing around R1 is absolutely wild to see. Full scale propaganda war.

Ironic
Post reply on HN