Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

161–170 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#161

I don’t think this entirely invalidates massive GPU spend just yet: “ Therefore, we can draw two conclusions: First, distilling more powerful models into smaller ones yields excellent results, whereas smaller models relying on the large-scale RL mentioned in this paper require enormous computational power and may not even achieve the performance of distillation. Second, while distillation strategies are both economic…

It does if the spend drives GPU prices so high that more researchers can't afford to use them. And DS demonstrated what a small team of researchers can do with a moderate amount of GPUs.

The DS team themselves suggest large amounts of compute are still required

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#162
post #44
post #9

we've been tracking the deepseek threads extensively in LS. related reads: - i consider the deepseek v3 paper required preread https://github.com/deepseek-ai/DeepSeek-V3 - R1 + Sonnet > R1 or O1 or R1+R1 or O1+Sonnet or any other combo https://aider.chat/2025/01/24/r1-sonnet.html - independent repros: 1) https://hkust-nlp.notion.site/simplerl-reason 2) https://buttondown.com/ainews/archive/ainews-tinyzero-reprod... 3…

I am extremely interested in your spam. Will you post it to https://www.latent.space/ ?

idk haha most of it is just twitter bookmarks - i will if i get to interview the deepseek team at some point (someone help put us in touch pls! swyx at ai.engineer )

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#164
post #151

Aside from the usual Tiananmen Square censorship, there's also some other propaganda baked-in: https://prnt.sc/HaSc4XZ89skA (from reddit)

Apparently the censorship isn't baked-in to the model itself, but rather is overlayed in the public chat interface. If you run it yourself, it is significantly less censored [0]

[0] https://thezvi.substack.com/p/on-deepseeks-r1?open=false#%C2...

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#165
post #151

Aside from the usual Tiananmen Square censorship, there's also some other propaganda baked-in: https://prnt.sc/HaSc4XZ89skA (from reddit)

In Communist theoretical texts the term "propaganda" is not negative and Communists are encouraged to produce propaganda to keep up morale in their own ranks and to produce propaganda that demoralize opponents.

The recent wave of the average Chinese has a better quality of life than the average Westerner propaganda is an obvious example of propaganda aimed at opponents.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#166
post #125

"Reasoning" will be disproven for this again within a few days I guess. Context: o1 does not reason, it pattern matches. If you rename variables, suddenly it fails to solve the request.

Perhaps, but over enough data pattern matching can becomes generalization ...

One of the interesting DeepSeek-R results is using a 1st generation (RL-trained) reasoning model to generate synthetic data (reasoning traces) to train a subsequent one, or even "distill" into a smaller model (by fine tuning the smaller model on this reasoning data).

Maybe "Data is all you need" (well, up to a point) ?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#168

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

Why do americans think china is like a hivemind controlled by an omnisicient Xi, making strategic moves to undermine them? Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x?

Because that’s the way China presents itself and that’s the way China boosters talk about China.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#169
post #119

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

Seeing what china is doing to the car market, I give it 5 years for China to do to the AI/GPU market to do the same. This will be good. Nvidia/OpenAI monopoly is bad for everyone. More competition will be welcome.

That is not going to happen without currently embargo'ed litography tech. They'd be already making more powerful GPUs if they could right now.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#170

Over 100 authors on arxiv and published under the team name, that's how you recognize everyone and build comradery. I bet morale is high over there

Same thing happened to Google Gemini paper (1000+ authors) and it was described as big co promo culture (everyone wants credits). Interesting how narratives shift

https://arxiv.org/abs/2403.05530

Post reply on HN