Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

11–20 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#11

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

How likely is this?

Just a cursory probing of deepseek yields all kinds of censoring of topics. Isn't it just as likely Chinese sponsors of this have incentivized and sponsored an undercutting of prices so that a more favorable LLM is preferred on the market?

Think about it, this is something they are willing to do with other industries.

And, if LLMs are going to be engineering accelerators as the world believes, then it wouldn't do to have your software assistants be built with a history book they didn't write. Better to dramatically subsidize your own domestic one then undercut your way to dominance.

It just so happens deepseek is the best one, but whichever was the best Chinese sponsored LLM would be the one we're supposed to use.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#12
post #4

I wonder if the decision to make o3-mini available for free user in the near (hopefully) future is a response to this really good, cheap and open reasoning model.

almost certainly (see chart) https://www.latent.space/p/reasoning-price-war (disclaimer i made it)

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#13

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

I've been confused over this. I've seen a $5.5M # for training, and commensurate commentary along the lines of what you said, but it elides the cost of the base model AFAICT.

$5.5 million is the cost of training the base model, DeepSeek V3. I haven't seen numbers for how much extra the reinforcement learning that turned it into R1 cost.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#14

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

More effecient use of hardware just increases productivity. No more people/teams can interate faster and in parralel

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#15

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

How likely is this? Just a cursory probing of deepseek yields all kinds of censoring of topics. Isn't it just as likely Chinese sponsors of this have incentivized and sponsored an undercutting of prices so that a more favorable LLM is preferred on the market? Think about it, this is something they are willing to do with other industries. And, if LLMs are going to be engineering accelerators as the world believes, the…

You raise an interesting point, and both of your points seem well-founded and have wide cache. However, I strongly believe both points are in error.

- OP elides costs of anything at all outside renting GPUs, and they purchased them, paid GPT-4 to generate training data, etc. etc.

- Non-Qwen models they trained are happy to talk about ex. Tiananmen

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#16

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

I've been confused over this. I've seen a $5.5M # for training, and commensurate commentary along the lines of what you said, but it elides the cost of the base model AFAICT.

With $5.5M, you can buy around 150 H100s. Experts correct me if I’m wrong but it’s practically impossible to train a model like that with that measly amount.

So I doubt that figure includes all the cost of training.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#17

Earlier quoted context omitted.

I've been confused over this. I've seen a $5.5M # for training, and commensurate commentary along the lines of what you said, but it elides the cost of the base model AFAICT.

$5.5 million is the cost of training the base model, DeepSeek V3. I haven't seen numbers for how much extra the reinforcement learning that turned it into R1 cost.

Ahhh, ty ty.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#19

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

But do we know that the same techniques won't scale if trained in the huge clusters?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#20

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

From what I've read, DeepSeek is a "side project" at a Chinese quant fund. They had the GPU capacity to spare.
Post reply on HN