https://prnt.sc/HaSc4XZ89skA (from reddit)
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
151–160 of 1001 posts
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#152https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#153"Reasoning" will be disproven for this again within a few days I guess. Context: o1 does not reason, it pattern matches. If you rename variables, suddenly it fails to solve the request.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#154I don’t think this entirely invalidates massive GPU spend just yet: “ Therefore, we can draw two conclusions: First, distilling more powerful models into smaller ones yields excellent results, whereas smaller models relying on the large-scale RL mentioned in this paper require enormous computational power and may not even achieve the performance of distillation. Second, while distillation strategies are both economic…
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#155I've always been leery about outrageous GPU investments, at some point I'll dig through and find my prior comments where I've said as much to that effect. The CEOs, upper management, and governments derive their importance on how much money they can spend - AI gave them the opportunity for them to confidently say that if you give me $X I can deliver Y and they turn around and give that money to NVidia. The problem wa…
Agree. The "need to build new buildings, new power plants, buy huge numbers of today's chips from one vendor" never made any sense considering we don't know what would be done in those buildings in 5 years when they're ready.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#156Larry Ellison is 80. Masayoshi Son is 67. Both have said that anti-aging and eternal life is one of their main goals with investing toward ASI. For them it's worth it to use their own wealth and rally the industry to invest $500 billion in GPUs if that means they will get to ASI 5 years faster and ask the ASI to give them eternal life.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#157I'm impressed by not only how good deepseek r1 is, but also how good the smaller distillations are. qwen-based 7b distillation of deepseek r1 is a great model too. the 32b distillation just became the default model for my home server.
How much VRAM is needed for the 32B distillation?
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#158I'm impressed by not only how good deepseek r1 is, but also how good the smaller distillations are. qwen-based 7b distillation of deepseek r1 is a great model too. the 32b distillation just became the default model for my home server.
How much VRAM is needed for the 32B distillation?
In general, if you're using 8bit which is virtually lossless, any dense model will require roughly the same amount as the number of params w/ a small context, and a bit more as you increase context.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#159I've always been leery about outrageous GPU investments, at some point I'll dig through and find my prior comments where I've said as much to that effect. The CEOs, upper management, and governments derive their importance on how much money they can spend - AI gave them the opportunity for them to confidently say that if you give me $X I can deliver Y and they turn around and give that money to NVidia. The problem wa…
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#160we've been tracking the deepseek threads extensively in LS. related reads: - i consider the deepseek v3 paper required preread https://github.com/deepseek-ai/DeepSeek-V3 - R1 + Sonnet > R1 or O1 or R1+R1 or O1+Sonnet or any other combo https://aider.chat/2025/01/24/r1-sonnet.html - independent repros: 1) https://hkust-nlp.notion.site/simplerl-reason 2) https://buttondown.com/ainews/archive/ainews-tinyzero-reprod... 3…
Hugging Face is reproducing R1 in public. https://x.com/_lewtun/status/1883142636820676965 https://github.com/huggingface/open-r1 Hugging Face Journal Club - DeepSeek R1 https://www.youtube.com/watch?v=1xDVbu-WaFo