Question about the rule-based rewards (correctness and format) mentioned in the paper: Does the raw base model just expected “stumble upon“ a correct answer /correct format to get a reward and start the learning process? Are there any more details about the reward modelling?
Good question. When BF Skinner used to train his pigeons, he’d initially reinforce any tiny movement that at least went in the right direction. For the exact reasons you mentioned. For example, instead of waiting for the pigeon to peck the lever directly (which it might not do for many hours), he’d give reinforcement if the pigeon so much as turned its head towards the lever. Over time, he’d raise the bar. Until, eve…
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
191–200 of 1001 posts
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#192The true costs and implications of V3 are discussed here: https://www.interconnects.ai/p/deepseek-v3-and-the-actual-co...
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#193I'm impressed by not only how good deepseek r1 is, but also how good the smaller distillations are. qwen-based 7b distillation of deepseek r1 is a great model too. the 32b distillation just became the default model for my home server.
How much VRAM is needed for the 32B distillation?
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#194The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.
Why do americans think china is like a hivemind controlled by an omnisicient Xi, making strategic moves to undermine them? Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x?
https://www.chinalawtranslate.com/en/generative-ai-interim/
In the case of TikTok, ByteDance and the government found ways to force international workers in the US to signing agreements that mirror local laws in mainland China:
https://dailycaller.com/2025/01/14/tiktok-forced-staff-oaths...
I find that degree of control to be dystopian and horrifying but I suppose it has helped their country focus and grow instead of dealing with internal conflict.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#195Earlier quoted context omitted.
have you tried asking chatgpt something even slightly controversial? chatgpt censors much more than deepseek does. also deepseek is open-weights. there is nothing preventing you from doing a finetune that removes the censorship. they did that with llama2 back in the day.
> chatgpt censors much more than deepseek does This is an outrageous claim with no evidence, as if there was any equivalence between government enforced propaganda and anything else. Look at the system prompts for DeepSeek and it’s even more clear. Also: fine tuning is not relevant when what is deployed at scale brainwashes the masses through false and misleading responses.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#196Earlier quoted context omitted.
Seeing what china is doing to the car market, I give it 5 years for China to do to the AI/GPU market to do the same. This will be good. Nvidia/OpenAI monopoly is bad for everyone. More competition will be welcome.
That is not going to happen without currently embargo'ed litography tech. They'd be already making more powerful GPUs if they could right now.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#197Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#198DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...
Meta is in full panic last I heard. They have amassed a collection of pseudo experts there to collect their checks. Yet, Zuck wants to keep burning money on mediocrity. I’ve yet to see anything of value in terms products out of Meta.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#199Earlier quoted context omitted.
The amount of astroturfing around R1 is absolutely wild to see. Full scale propaganda war.
I would argue there is too little hype given the downloadable models for Deep Seek. There should be alot of hype around this organically. If anything, the other half good fully closed non ChatGPT models are astroturfing. I made a post in december 2023 whining about the non hype for Deep Seek. https://news.ycombinator.com/item?id=38505986
There’s a lot of astroturfing from a lot of different parties for a few different reasons. Which is all very interesting.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#200Earlier quoted context omitted.
Meta is in full panic last I heard. They have amassed a collection of pseudo experts there to collect their checks. Yet, Zuck wants to keep burning money on mediocrity. I’ve yet to see anything of value in terms products out of Meta.
I guess all that leetcoding and stack ranking didn't in fact produce "the cream of the crop"...