Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

181–190 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#181

I've always been leery about outrageous GPU investments, at some point I'll dig through and find my prior comments where I've said as much to that effect. The CEOs, upper management, and governments derive their importance on how much money they can spend - AI gave them the opportunity for them to confidently say that if you give me $X I can deliver Y and they turn around and give that money to NVidia. The problem wa…

IMO the you cannot fail by investing in compute. If it turns out you only need 1/1000th of the compute to train and or run your models, great! Now you can spend that compute on inference that solves actual problems humans have.

o3 $4k compute spend per task made it pretty clear that once we reach AGI inference is going to be the majority of spend. We'll spend compute getting AI to cure cancer or improve itself rather than just training at chatbot that helps students cheat on their exams. The more compute you have, the more problems you can solve faster, the bigger your advantage, especially if/when recursive self improvement kicks off, efficiency improvements only widen this gap

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#182

DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...

Correct me if I'm wrong but if Chinese can produce the same quality at %99 discount, then the supposed $500B investment is actually worth $5B. Isn't that the kind wrong investment that can break nations?

Edit: Just to clarify, I don't imply that this is public money to be spent. It will commission $500B worth of human and material resources for 5 years that can be much more productive if used for something else - i.e. high speed rail network instead of a machine that Chinese built for $5B.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#183
post #167

Over 100 authors on arxiv and published under the team name, that's how you recognize everyone and build comradery. I bet morale is high over there

It’s credential stuffing.

Come on man, let them have their well deserved win as a team.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#184
post #151

Aside from the usual Tiananmen Square censorship, there's also some other propaganda baked-in: https://prnt.sc/HaSc4XZ89skA (from reddit)

Apparently the censorship isn't baked-in to the model itself, but rather is overlayed in the public chat interface. If you run it yourself, it is significantly less censored [0] [0] https://thezvi.substack.com/p/on-deepseeks-r1?open=false#%C2...

Interestingly they cite for the Tiananmen Square prompt a Tweet[1] that shows the poster used the Distilled Llama model, which per a reply Tweet (quoted below) doesn't transfer the safety/censorship layer. While others using the non-Distilled model encounter the censorship when locally hosted.

> You're running Llama-distilled R1 locally. Distillation transfers the reasoning process, but not the "safety" post-training. So you see the answer mostly from Llama itself. R1 refuses to answer this question without any system prompt (official API or locally).

[1] https://x.com/PerceivingAI/status/1881504959306273009

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#186
post #151

Aside from the usual Tiananmen Square censorship, there's also some other propaganda baked-in: https://prnt.sc/HaSc4XZ89skA (from reddit)

At least it’s not home grown propaganda from the US, so will likely not cover most other topics of interest.

What are you basing this whataboutism on?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#187
post #31

Earlier quoted context omitted.

Alexandr Wang did not even say they lied in the paper. Here's the interview: https://www.youtube.com/watch?v=x9Ekl9Izd38 . "My understanding is that is that Deepseek has about 50000 a100s, which they can't talk about obviously, because it is against the export controls that the United States has put in place. And I think it is true that, you know, I think they have more chips than other people expect..." Plus, how ex…

> Plus, how exactly did Deepseek lie. The model size, data size are all known. Calculating the number of FLOPS is an exercise in arithmetics, which is perhaps the secret Deepseek has because it seemingly eludes people. Model parameter count and training set token count are fixed. But other things such as epochs are not. In the same amount of time, you could have 1 epoch or 100 epochs depending on how many GPUs you ha…

> In the same amount of time, you could have 1 epoch or 100 epochs depending on how many GPUs you have.

This is just not true for RL and related algorithms, having more GPU/agents encounters diminishing returns, and is just not the equivalent to letting a single agent go through more steps.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#188
post #131

I'm impressed by not only how good deepseek r1 is, but also how good the smaller distillations are. qwen-based 7b distillation of deepseek r1 is a great model too. the 32b distillation just became the default model for my home server.

I just tries the distilled 8b Llama variant, and it had very poor prompt adherence.

It also reasoned its way to an incorrect answer, to a question plain Llama 3.1 8b got fairly correct.

So far not impressed, but will play with the qwen ones tomorrow.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#189
post #146

Larry Ellison is 80. Masayoshi Son is 67. Both have said that anti-aging and eternal life is one of their main goals with investing toward ASI. For them it's worth it to use their own wealth and rally the industry to invest $500 billion in GPUs if that means they will get to ASI 5 years faster and ask the ASI to give them eternal life.

Probably shouldn't be firing their blood boys just yet ... According to Musk, SoftBank only has $10B available for this atm.

Elon says a lot of things.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#190
post #121

Earlier quoted context omitted.

CEO of Scale said Deepseek is lying and actually has a 50k GPU cluster. He said they lied in the paper because technically they aren't supposed to have them due to export laws. I feel like this is very likely. They obvious did some great breakthroughs, but I doubt they were able to train on so much less hardware.

Why would Deepseek lie? They are in China, American export laws can't touch them.

Making it obvious that they managed to circumvent sanctions isn’t going to help them. It will turn public sentiment in the west even more against them and will motivate politicians to make the enforcement stricter and prevent GPU exports.
Post reply on HN