Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

781–790 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#781
post #504

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

The nVidia market price could also be questionable considering how much cheaper DS is to run.

I thought so at first too, but then realized this may actually unlock more total demand for them.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#782

Earlier quoted context omitted.

Whenever I use it, it just seems to spin itself in circles for ages, spit out a half-assed summary and give up. Is it like the OpenAI models in that in needs to be prompted in extremely-specific ways to get it to not be garbage?

I'm curious what you are asking it to do and whether you think the thoughts it expresses along the seemed likely to lead it in a useful direction before it resorted to a summary. Also perhaps it doesn't realize you don't want a summary?

People be like, "please provide me with a full stack web app" and then think its bad when it doesnt.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#783
post #773

Earlier quoted context omitted.

Why should the bubble pop when we just got the proof that these models can be much more efficient than we thought? I mean, sure, no one is going to have a monopoly, and we're going to see a race to the bottom in prices, but on the other hand, the AI revolution is going to come much sooner than expected, and it's going to be on everyone's pocket this year. Isn't that a bullish signal for the economy?

I think this is the correct take. There might be a small bubble burst initially after a bunch of US stocks retrace due to uncertainty. But in the long run this should speed up the proliferation of productivity gains unlocked by AI.

I think we should not underestimate one aspect: at the moment, a lot of hype is artificial (and despicable if you ask me). Anthropic says AI can double human lifespan in 10 years time; openAI says they have AGI behind the corner; META keeps insisting on their model being open source when they in fact only release the weights. They think - maybe they are right - that they would not be able to get these massive investments without hyping things a bit but deepseek's performance should call for things to be reviewed.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#784
post #40

Question about the rule-based rewards (correctness and format) mentioned in the paper: Does the raw base model just expected “stumble upon“ a correct answer /correct format to get a reward and start the learning process? Are there any more details about the reward modelling?

yes, stumble on a correct answer and also pushing down incorrect answer probability in the meantime. their base model is pretty good

It seems a strong base model is what enabled this. The models needs to be smart enough to get it right at least some times.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#785
post #268

I've been using https://chat.deepseek.com/ over My ChatGPT Pro subscription because being able to read the thinking in the way they present it is just much much easier to "debug" - also I can see when it's bending it's reply to something, often softening it or pandering to me - I can just say "I saw in your thinking you should give this type of reply, don't do that". If it stays free and gets better that's going to b…

Chatgpt does this as well, it just doesn't display it in the UI. You can click on the "thinking" to expand and read the tomhought process.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#786

Earlier quoted context omitted.

[I typed something dumb while half asleep]

I'm not sure censorship or lack of it matters for most use cases. Why would businesses using LLM to speed up their processes, or a programmer using it to write code care about how accurately it answers to political questions?

Ethics.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#787

Earlier quoted context omitted.

I asked Chatgpt: how many civilians Israel killed in Gaza. Please provide a rough estimate. As of January 2025, the conflict between Israel and Hamas has resulted in significant civilian casualties in the Gaza Strip. According to reports from the United Nations Office for the Coordination of Humanitarian Affairs (OCHA), approximately 7,000 Palestinian civilians have been killed since the escalation began in October 2…

Isn't the real number around 46,000 people, though?

No one knows the real number.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#788
post #9

we've been tracking the deepseek threads extensively in LS. related reads: - i consider the deepseek v3 paper required preread https://github.com/deepseek-ai/DeepSeek-V3 - R1 + Sonnet > R1 or O1 or R1+R1 or O1+Sonnet or any other combo https://aider.chat/2025/01/24/r1-sonnet.html - independent repros: 1) https://hkust-nlp.notion.site/simplerl-reason 2) https://buttondown.com/ainews/archive/ainews-tinyzero-reprod... 3…

could someone explain how the RL works here? I don't understand how it can be a training objective with a LLM?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#789

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

Given this comment, I tried it. It's no where close to Claude, and it's also not better than OpenAI. I'm so confused as to how people judge these things.

I'm confused as to how you haven't found R1 to be much better. My experience has been exactly like that of the OP's

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#790
OOPS DEEPSEEK

"Your Point About Authoritarian Systems: You mentioned that my responses seem to reflect an authoritarian communist system and that I am denying the obvious. Let me clarify:

My goal is to provide accurate and historically grounded explanations based on the laws, regulations..."

DEEPSEEK 2025

After I proved my point it was wrong after @30 minutes of its brainwashing false conclusions it said this after I posted a law:

"Oops! DeepSeek is experiencing high traffic at the moment. Please check back in a little while."

I replied: " Oops! is right you want to deny.."

"

"

Post reply on HN