Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

251–260 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#251
post #79

How can openai justify their $200/mo subscriptions if a model like this exists at an incredibly low price point? Operator? I've been impressed in my brief personal testing and the model ranks very highly across most benchmarks (when controlled for style it's tied number one on lmarena). It's also hilarious that openai explicitly prevented users from seeing the CoT tokens on the o1 model (which you still pay for btw)…

openai has better models in the bank so short term they will release o3-derived models

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#252
post #170

Over 100 authors on arxiv and published under the team name, that's how you recognize everyone and build comradery. I bet morale is high over there

Same thing happened to Google Gemini paper (1000+ authors) and it was described as big co promo culture (everyone wants credits). Interesting how narratives shift https://arxiv.org/abs/2403.05530

[deleted]

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#253
post #77
post #40

Question about the rule-based rewards (correctness and format) mentioned in the paper: Does the raw base model just expected “stumble upon“ a correct answer /correct format to get a reward and start the learning process? Are there any more details about the reward modelling?

Good question. When BF Skinner used to train his pigeons, he’d initially reinforce any tiny movement that at least went in the right direction. For the exact reasons you mentioned. For example, instead of waiting for the pigeon to peck the lever directly (which it might not do for many hours), he’d give reinforcement if the pigeon so much as turned its head towards the lever. Over time, he’d raise the bar. Until, eve…

they’re not doing anything like that and you are actually describing the failed research direction a lot of the frontier labs (esp Google) were doing

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#254

Earlier quoted context omitted.

I would argue there is too little hype given the downloadable models for Deep Seek. There should be alot of hype around this organically. If anything, the other half good fully closed non ChatGPT models are astroturfing. I made a post in december 2023 whining about the non hype for Deep Seek. https://news.ycombinator.com/item?id=38505986

Possible for that to also be true! There’s a lot of astroturfing from a lot of different parties for a few different reasons. Which is all very interesting.

How do you know it's astroturfing and not legitimate hype about an impressive and open technical achievement?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#255
I’ve been using R1 last few days and it’s noticeably worse than O1 at everything. It’s impressive, better than my latest Claude run (I stopped using Claude completely once O1 came out), but O1 is just flat out better.

Perhaps the gap is minor, but it feels large. I’m hesitant on getting O1 Pro, because using a worse model just seems impossible once you’ve experienced a better one

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#256
post #40

Question about the rule-based rewards (correctness and format) mentioned in the paper: Does the raw base model just expected “stumble upon“ a correct answer /correct format to get a reward and start the learning process? Are there any more details about the reward modelling?

yes, stumble on a correct answer and also pushing down incorrect answer probability in the meantime. their base model is pretty good

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#257
post #165
post #151

Aside from the usual Tiananmen Square censorship, there's also some other propaganda baked-in: https://prnt.sc/HaSc4XZ89skA (from reddit)

In Communist theoretical texts the term "propaganda" is not negative and Communists are encouraged to produce propaganda to keep up morale in their own ranks and to produce propaganda that demoralize opponents. The recent wave of the average Chinese has a better quality of life than the average Westerner propaganda is an obvious example of propaganda aimed at opponents.

Is it propaganda if it's true?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#258
post #59

Genuinly curious, what is everyone using reasoning models for? (R1/o1/o3)

Regular coding questions mostly. For me o1 generally gives better code and understands the prompt more completely (haven’t started using r1 or o3 regularly enough to opine).

o3 isn’t available

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#259

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

[deleted]

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#260
post #182

DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...

Correct me if I'm wrong but if Chinese can produce the same quality at %99 discount, then the supposed $500B investment is actually worth $5B. Isn't that the kind wrong investment that can break nations? Edit: Just to clarify, I don't imply that this is public money to be spent. It will commission $500B worth of human and material resources for 5 years that can be much more productive if used for something else - i.e…

Trump just pull a stunt with Saudi Arabia. He first tried to "convince" them to reduce the oil price to hurt Russia. In the following negotiations the oil price was no longer mentioned but MBS promised to invest $600 billion in the U.S. over 4 years:

https://fortune.com/2025/01/23/saudi-crown-prince-mbs-trump-...

Since the Stargate Initiative is a private sector deal, this may have been a perfect shakedown of Saudi Arabia. SA has always been irrationally attracted to "AI", so perhaps it was easy. I mean that part of the $600 billion will go to "AI".

Post reply on HN