Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

41–50 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#41

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

Why do americans think china is like a hivemind controlled by an omnisicient Xi, making strategic moves to undermine them? Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x?

> Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x

They literally published all their methodology. It's nothing groundbreaking, just western labs seem slow to adopt new research. Mixture of experts, key-value cache compression, multi-token prediction, 2/3 of these weren't invented by DeepSeek. They did invent a new hardware-aware distributed training approach for mixture-of-experts training that helped a lot, but there's nothing super genius about it, western labs just never even tried to adjust their model to fit the hardware available.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#43

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

[deleted]

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#44
post #9

we've been tracking the deepseek threads extensively in LS. related reads: - i consider the deepseek v3 paper required preread https://github.com/deepseek-ai/DeepSeek-V3 - R1 + Sonnet > R1 or O1 or R1+R1 or O1+Sonnet or any other combo https://aider.chat/2025/01/24/r1-sonnet.html - independent repros: 1) https://hkust-nlp.notion.site/simplerl-reason 2) https://buttondown.com/ainews/archive/ainews-tinyzero-reprod... 3…

I am extremely interested in your spam. Will you post it to https://www.latent.space/ ?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#46

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

CEO of Scale said Deepseek is lying and actually has a 50k GPU cluster. He said they lied in the paper because technically they aren't supposed to have them due to export laws. I feel like this is very likely. They obvious did some great breakthroughs, but I doubt they were able to train on so much less hardware.

I haven't had time to follow this thread, but it looks like some people are starting to experimentally replicate DeepSeek on extremely limited H100 training:

> You can RL post-train your small LLM (on simple tasks) with only 10 hours of H100s.

https://www.reddit.com/r/singularity/comments/1i99ebp/well_s...

Forgive me if this is inaccurate. I'm rushing around too much this afternoon to dive in.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#47

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

Why do americans think china is like a hivemind controlled by an omnisicient Xi, making strategic moves to undermine them? Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x?

China is actually just one person (Xi) acting in perfect unison and its purpose is not to benefit its own people, but solely to undermine the West.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#48
post #26

Earlier quoted context omitted.

CEO of Scale said Deepseek is lying and actually has a 50k GPU cluster. He said they lied in the paper because technically they aren't supposed to have them due to export laws. I feel like this is very likely. They obvious did some great breakthroughs, but I doubt they were able to train on so much less hardware.

I would think the CEO of an American AI company has every reason to neg and downplay foreign competition... And since it's a businessperson they're going to make it sound as cute and innocuous as possible

Or, more likely, there wasn't a magic innovation that nobody else thought of, that reduced costs by orders of magnitude.

When deciding between mostly like scenarios, it is more likely that the company lied than they found some industry changing magic innovation.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#49

Earlier quoted context omitted.

Why do americans think china is like a hivemind controlled by an omnisicient Xi, making strategic moves to undermine them? Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x?

> Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x They literally published all their methodology. It's nothing groundbreaking, just western labs seem slow to adopt new research. Mixture of experts, key-value cache compression, multi-token prediction, 2/3 of these weren't invented by DeepSeek. They did invent a new hardware-aware distributed training approach for mixture…

But those approaches alone wouldn’t yield the improvements claimed. How did they train the foundational model upon which they applied RL, distillations, etc? That part is unclear and I don’t think anything they’ve released anything that explains the low cost.

It’s also curious why some people are seeing responses where it thinks it is an OpenAI model. I can’t find the post but someone had shared a link to X with that in one of the other HN discussions.

Post reply on HN