DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
1–10 of 1001 posts
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#2Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#3I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#4Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#5The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#6The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.
I've seen a $5.5M # for training, and commensurate commentary along the lines of what you said, but it elides the cost of the base model AFAICT.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#7Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#8The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.
DeepSeek's R1 also blew all the other China LLM teams out of the water, in spite of their larger training budgets and greater hardware resources (e.g. Alibaba). I suspect it's because its creators' background in a trading firm made them more willing to take calculated risks and incorporate all the innovations that made R1 such a success, rather than just copying what other teams are doing with minimal innovation.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#9- i consider the deepseek v3 paper required preread https://github.com/deepseek-ai/DeepSeek-V3
- R1 + Sonnet > R1 or O1 or R1+R1 or O1+Sonnet or any other combo https://aider.chat/2025/01/24/r1-sonnet.html
- independent repros: 1) https://hkust-nlp.notion.site/simplerl-reason 2) https://buttondown.com/ainews/archive/ainews-tinyzero-reprod... 3) https://x.com/ClementDelangue/status/1883154611348910181
- R1 distillations are going to hit us every few days - because it's ridiculously easy (https://buttondown.com/ainews/archive/ainews-bespoke-stratos... , 23min interview w team https://www.youtube.com/watch?v=jrf76uNs77k)
i probably have more resources but dont want to spam - seek out the latent space discord if you want the full stream i pulled these notes from
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#10Over 100 authors on that paper. Cred stuffing ftw.