Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

1–10 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#3
The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value.

I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#5

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

The US economy is predicated on the perception that AI requires a lot of GPUs? That seems like a stretch.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#6

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

I've been confused over this.

I've seen a $5.5M # for training, and commensurate commentary along the lines of what you said, but it elides the cost of the base model AFAICT.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#8

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

>I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

DeepSeek's R1 also blew all the other China LLM teams out of the water, in spite of their larger training budgets and greater hardware resources (e.g. Alibaba). I suspect it's because its creators' background in a trading firm made them more willing to take calculated risks and incorporate all the innovations that made R1 such a success, rather than just copying what other teams are doing with minimal innovation.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#9
we've been tracking the deepseek threads extensively in LS. related reads:

- i consider the deepseek v3 paper required preread https://github.com/deepseek-ai/DeepSeek-V3

- R1 + Sonnet > R1 or O1 or R1+R1 or O1+Sonnet or any other combo https://aider.chat/2025/01/24/r1-sonnet.html

- independent repros: 1) https://hkust-nlp.notion.site/simplerl-reason 2) https://buttondown.com/ainews/archive/ainews-tinyzero-reprod... 3) https://x.com/ClementDelangue/status/1883154611348910181

- R1 distillations are going to hit us every few days - because it's ridiculously easy (https://buttondown.com/ainews/archive/ainews-bespoke-stratos... , 23min interview w team https://www.youtube.com/watch?v=jrf76uNs77k)

i probably have more resources but dont want to spam - seek out the latent space discord if you want the full stream i pulled these notes from

Post reply on HN