Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

191–200 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#191
post #77
post #40

Question about the rule-based rewards (correctness and format) mentioned in the paper: Does the raw base model just expected “stumble upon“ a correct answer /correct format to get a reward and start the learning process? Are there any more details about the reward modelling?

Good question. When BF Skinner used to train his pigeons, he’d initially reinforce any tiny movement that at least went in the right direction. For the exact reasons you mentioned. For example, instead of waiting for the pigeon to peck the lever directly (which it might not do for many hours), he’d give reinforcement if the pigeon so much as turned its head towards the lever. Over time, he’d raise the bar. Until, eve…

Since intermediate steps of reasoning are hard to verify they only award final results. Yet that produces enough signal to produce more productive reasoning over time. In a way when pigeons are virtual one can afford to have a lot more of them.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#192
For context: R1 is a reasoning model based on V3. DeepSeek has claimed that GPU costs to train V3 (given prevailing rents) were about $5M.

The true costs and implications of V3 are discussed here: https://www.interconnects.ai/p/deepseek-v3-and-the-actual-co...

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#193
post #131

I'm impressed by not only how good deepseek r1 is, but also how good the smaller distillations are. qwen-based 7b distillation of deepseek r1 is a great model too. the 32b distillation just became the default model for my home server.

How much VRAM is needed for the 32B distillation?

I had no problems running the 32b at q4 quantization with 24GB of ram.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#194

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

Why do americans think china is like a hivemind controlled by an omnisicient Xi, making strategic moves to undermine them? Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x?

Well it is like a hive mind due to the degree of control. Most Chinese companies are required by law to literally uphold the country’s goals - see translation of Chinese law, which says generative AI must uphold their socialist values:

https://www.chinalawtranslate.com/en/generative-ai-interim/

In the case of TikTok, ByteDance and the government found ways to force international workers in the US to signing agreements that mirror local laws in mainland China:

https://dailycaller.com/2025/01/14/tiktok-forced-staff-oaths...

I find that degree of control to be dystopian and horrifying but I suppose it has helped their country focus and grow instead of dealing with internal conflict.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#195
post #115

Earlier quoted context omitted.

have you tried asking chatgpt something even slightly controversial? chatgpt censors much more than deepseek does. also deepseek is open-weights. there is nothing preventing you from doing a finetune that removes the censorship. they did that with llama2 back in the day.

> chatgpt censors much more than deepseek does This is an outrageous claim with no evidence, as if there was any equivalence between government enforced propaganda and anything else. Look at the system prompts for DeepSeek and it’s even more clear. Also: fine tuning is not relevant when what is deployed at scale brainwashes the masses through false and misleading responses.

[flagged]

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#196
post #119

Earlier quoted context omitted.

Seeing what china is doing to the car market, I give it 5 years for China to do to the AI/GPU market to do the same. This will be good. Nvidia/OpenAI monopoly is bad for everyone. More competition will be welcome.

That is not going to happen without currently embargo'ed litography tech. They'd be already making more powerful GPUs if they could right now.

they seem to be doing fine so far. every day we wake up to more success stories from china's AI/semiconductory industry.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#197

Earlier quoted context omitted.

The amount of astroturfing around R1 is absolutely wild to see. Full scale propaganda war.

Ironic

That word does not mean what you think it means.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#198
post #175

DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...

Meta is in full panic last I heard. They have amassed a collection of pseudo experts there to collect their checks. Yet, Zuck wants to keep burning money on mediocrity. I’ve yet to see anything of value in terms products out of Meta.

I guess all that leetcoding and stack ranking didn't in fact produce "the cream of the crop"...

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#199

Earlier quoted context omitted.

The amount of astroturfing around R1 is absolutely wild to see. Full scale propaganda war.

I would argue there is too little hype given the downloadable models for Deep Seek. There should be alot of hype around this organically. If anything, the other half good fully closed non ChatGPT models are astroturfing. I made a post in december 2023 whining about the non hype for Deep Seek. https://news.ycombinator.com/item?id=38505986

Possible for that to also be true!

There’s a lot of astroturfing from a lot of different parties for a few different reasons. Which is all very interesting.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#200
post #175

Earlier quoted context omitted.

Meta is in full panic last I heard. They have amassed a collection of pseudo experts there to collect their checks. Yet, Zuck wants to keep burning money on mediocrity. I’ve yet to see anything of value in terms products out of Meta.

I guess all that leetcoding and stack ranking didn't in fact produce "the cream of the crop"...

It produces the cream of the leetcoding stack ranking crop.
Post reply on HN