I don’t think this entirely invalidates massive GPU spend just yet: “ Therefore, we can draw two conclusions: First, distilling more powerful models into smaller ones yields excellent results, whereas smaller models relying on the large-scale RL mentioned in this paper require enormous computational power and may not even achieve the performance of distillation. Second, while distillation strategies are both economic…
It does if the spend drives GPU prices so high that more researchers can't afford to use them. And DS demonstrated what a small team of researchers can do with a moderate amount of GPUs.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
161–170 of 1001 posts
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#162we've been tracking the deepseek threads extensively in LS. related reads: - i consider the deepseek v3 paper required preread https://github.com/deepseek-ai/DeepSeek-V3 - R1 + Sonnet > R1 or O1 or R1+R1 or O1+Sonnet or any other combo https://aider.chat/2025/01/24/r1-sonnet.html - independent repros: 1) https://hkust-nlp.notion.site/simplerl-reason 2) https://buttondown.com/ainews/archive/ainews-tinyzero-reprod... 3…
I am extremely interested in your spam. Will you post it to https://www.latent.space/ ?
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#163Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#164Aside from the usual Tiananmen Square censorship, there's also some other propaganda baked-in: https://prnt.sc/HaSc4XZ89skA (from reddit)
[0] https://thezvi.substack.com/p/on-deepseeks-r1?open=false#%C2...
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#165Aside from the usual Tiananmen Square censorship, there's also some other propaganda baked-in: https://prnt.sc/HaSc4XZ89skA (from reddit)
The recent wave of the average Chinese has a better quality of life than the average Westerner propaganda is an obvious example of propaganda aimed at opponents.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#166"Reasoning" will be disproven for this again within a few days I guess. Context: o1 does not reason, it pattern matches. If you rename variables, suddenly it fails to solve the request.
One of the interesting DeepSeek-R results is using a 1st generation (RL-trained) reasoning model to generate synthetic data (reasoning traces) to train a subsequent one, or even "distill" into a smaller model (by fine tuning the smaller model on this reasoning data).
Maybe "Data is all you need" (well, up to a point) ?
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#167Over 100 authors on arxiv and published under the team name, that's how you recognize everyone and build comradery. I bet morale is high over there
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#168The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.
Why do americans think china is like a hivemind controlled by an omnisicient Xi, making strategic moves to undermine them? Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x?
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#169The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.
Seeing what china is doing to the car market, I give it 5 years for China to do to the AI/GPU market to do the same. This will be good. Nvidia/OpenAI monopoly is bad for everyone. More competition will be welcome.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#170Over 100 authors on arxiv and published under the team name, that's how you recognize everyone and build comradery. I bet morale is high over there