Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

231–240 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#231
post #175

Earlier quoted context omitted.

Meta is in full panic last I heard. They have amassed a collection of pseudo experts there to collect their checks. Yet, Zuck wants to keep burning money on mediocrity. I’ve yet to see anything of value in terms products out of Meta.

DeepSeek was built on the foundations of public research, a major part of which is the Llama family of models. Prior to Llama open weights LLMs were considerably less performant; without Llama we might not have gotten Mistral, Qwen, or DeepSeek. This isn't meant to diminish DeepSeek's contributions, however: they've been doing great work on mixture of experts models and really pushing the community forward on that fr…

As far as I know, Llama's architecture has always been quite conservative: it has not changed that much since LLaMA. Most of their recent gains have been in post-training.

That's not to say their work is unimpressive or not worthy - as you say, they've facilitated much of the open-source ecosystem and have been an enabling factor for many - but it's more that that work has been in making it accessible, not necessarily pushing the frontier of what's actually possible, and DeepSeek has shown us what's possible when you do the latter.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#232
post #205

Earlier quoted context omitted.

I don't say that at all. Money spent on BS still sucks resources, no matter who spends that money. They are not going to make the GPU's from 500 billion dollar banknotes, they will pay people $500B to work on this stuff which means people won't be working on other stuff that can actually produce value worth more than the $500B. I guess the power plants are salvageable.

By that logic all money is waste. The money isnt destroyed when it is spent. It is transferred into someone else's bank account only. This process repeats recursively until taxation returns all money back to the treasury to be spent again. And out of this process of money shuffling: entire nations full of power plants!

Money is just IOUs, it means for some reason not specified on the banknote you are owed services. If in a society a small group of people are owed all the services they can indeed commission all those people.

If your rich spend all their money on building pyramids you end up with pyramids instead of something else. They could have chosen to make irrigation systems and have a productive output that makes the whole society more prosperous. Either way the workers get their money, on the Pyramid option their money ends up buying much less food though.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#233
post #203
post #182

Earlier quoted context omitted.

Correct me if I'm wrong but if Chinese can produce the same quality at %99 discount, then the supposed $500B investment is actually worth $5B. Isn't that the kind wrong investment that can break nations? Edit: Just to clarify, I don't imply that this is public money to be spent. It will commission $500B worth of human and material resources for 5 years that can be much more productive if used for something else - i.e…

Sigh, I don't understand why they had to do the $500 billion announcement with the president. So many people now wrongly think Trump just gave OpenAI $500 billion of the taxpayers' money.

It means he’ll knock down regulatory barriers and mess with competitors because his brand is associated with it. It was a smart poltical move by OpenAI.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#234
post #175

Earlier quoted context omitted.

Meta is in full panic last I heard. They have amassed a collection of pseudo experts there to collect their checks. Yet, Zuck wants to keep burning money on mediocrity. I’ve yet to see anything of value in terms products out of Meta.

I guess all that leetcoding and stack ranking didn't in fact produce "the cream of the crop"...

You sound extremely satisfied by that. I'm glad you found a way to validate your preconceived notions on this beautiful day. I hope your joy is enduring.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#235
post #182

DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...

Correct me if I'm wrong but if Chinese can produce the same quality at %99 discount, then the supposed $500B investment is actually worth $5B. Isn't that the kind wrong investment that can break nations? Edit: Just to clarify, I don't imply that this is public money to be spent. It will commission $500B worth of human and material resources for 5 years that can be much more productive if used for something else - i.e…

The $500B is just an aspirational figure they hope to spend on data centers to run AI models, such as GPT-o1 and its successors, that have already been developed.

If you want to compare the DeepSeek-R development costs to anything, you should be comparing it to what it cost OpenAI to develop GPT-o1 (not what they plan to spend to run it), but both numbers are somewhat irrelevant since they both build upon prior research.

Perhaps what's more relevant is that DeepSeek are not only open sourcing DeepSeek-R1, but have described in a fair bit of detail how they trained it, and how it's possible to use data generated by such a model to fine-tune a much smaller model (without needing RL) to much improve it's "reasoning" performance.

This is all raising the bar on the performance you can get for free, or run locally, which reduces what companies like OpenAI can charge for it.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#236
post #112
post #53

Earlier quoted context omitted.

You can try it yourself, it's refreshingly good.

Agreed. I am no fan of the CCP but I have no issue with using DeepSeek since I only need to use it for coding which it does quite well. I still believe Sonnet is better. DeepSeek also struggles when the context window gets big. This might be hardware though. Having said that, DeepSeek is 10 times cheaper than Sonnet and better than GPT-4o for my use cases. Models are a commodity product and it is easy enough to add a…

Curious why you have to qualify this with a “no fan of the CCP” prefix. From the outset, this is just a private organization and its links to CCP aren’t any different than, say, Foxconn’s or DJI’s or any of the countless Chinese manufacturers and businesses

You don’t invoke “I’m no fan of the CCP” before opening TikTok or buying a DJI drone or a BYD car. Then why this, because I’ve seen the same line repeated everywhere

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#237
post #131

I'm impressed by not only how good deepseek r1 is, but also how good the smaller distillations are. qwen-based 7b distillation of deepseek r1 is a great model too. the 32b distillation just became the default model for my home server.

Great as long as you’re not interested in Tiananmen Square or the Uighurs.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#238
post #35

Reddit's /r/chatgpt subreddit is currently heavily brigaded by bots/shills praising r1, I'd be very suspicious of any claims about it.

The amount of astroturfing around R1 is absolutely wild to see. Full scale propaganda war.

The literal creator of Netscape Navigator is going ga-ga over it on Twitter and HN thinks its all botted

This is not a serious place

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#239

For context: R1 is a reasoning model based on V3. DeepSeek has claimed that GPU costs to train V3 (given prevailing rents) were about $5M. The true costs and implications of V3 are discussed here: https://www.interconnects.ai/p/deepseek-v3-and-the-actual-co...

Thank you for providing this context and sourcing. I've been trying to find the root and details around the $5 million claim

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#240
post #211
post #175

Earlier quoted context omitted.

Meta is in full panic last I heard. They have amassed a collection of pseudo experts there to collect their checks. Yet, Zuck wants to keep burning money on mediocrity. I’ve yet to see anything of value in terms products out of Meta.

I would think Meta - who open source their model - would be less freaked out than those others that do not.

The criticism seems to mostly be that Meta maintains very expensive cost structure and fat organisation in the AI. While Meta can afford to do this, if smaller orgs can produce better results it means Meta is paying a lot for nothing. Meta shareholders now need to ask the question how many non-productive people Meta is employing and is Zuck in the control of the cost.
Post reply on HN