For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…
The nVidia market price could also be questionable considering how much cheaper DS is to run.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
781–790 of 1001 posts
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#782Earlier quoted context omitted.
Whenever I use it, it just seems to spin itself in circles for ages, spit out a half-assed summary and give up. Is it like the OpenAI models in that in needs to be prompted in extremely-specific ways to get it to not be garbage?
I'm curious what you are asking it to do and whether you think the thoughts it expresses along the seemed likely to lead it in a useful direction before it resorted to a summary. Also perhaps it doesn't realize you don't want a summary?
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#783Earlier quoted context omitted.
Why should the bubble pop when we just got the proof that these models can be much more efficient than we thought? I mean, sure, no one is going to have a monopoly, and we're going to see a race to the bottom in prices, but on the other hand, the AI revolution is going to come much sooner than expected, and it's going to be on everyone's pocket this year. Isn't that a bullish signal for the economy?
I think this is the correct take. There might be a small bubble burst initially after a bunch of US stocks retrace due to uncertainty. But in the long run this should speed up the proliferation of productivity gains unlocked by AI.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#784Question about the rule-based rewards (correctness and format) mentioned in the paper: Does the raw base model just expected “stumble upon“ a correct answer /correct format to get a reward and start the learning process? Are there any more details about the reward modelling?
yes, stumble on a correct answer and also pushing down incorrect answer probability in the meantime. their base model is pretty good
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#785I've been using https://chat.deepseek.com/ over My ChatGPT Pro subscription because being able to read the thinking in the way they present it is just much much easier to "debug" - also I can see when it's bending it's reply to something, often softening it or pandering to me - I can just say "I saw in your thinking you should give this type of reply, don't do that". If it stays free and gets better that's going to b…
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#786Earlier quoted context omitted.
[I typed something dumb while half asleep]
I'm not sure censorship or lack of it matters for most use cases. Why would businesses using LLM to speed up their processes, or a programmer using it to write code care about how accurately it answers to political questions?
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#787Earlier quoted context omitted.
I asked Chatgpt: how many civilians Israel killed in Gaza. Please provide a rough estimate. As of January 2025, the conflict between Israel and Hamas has resulted in significant civilian casualties in the Gaza Strip. According to reports from the United Nations Office for the Coordination of Humanitarian Affairs (OCHA), approximately 7,000 Palestinian civilians have been killed since the escalation began in October 2…
Isn't the real number around 46,000 people, though?
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#788we've been tracking the deepseek threads extensively in LS. related reads: - i consider the deepseek v3 paper required preread https://github.com/deepseek-ai/DeepSeek-V3 - R1 + Sonnet > R1 or O1 or R1+R1 or O1+Sonnet or any other combo https://aider.chat/2025/01/24/r1-sonnet.html - independent repros: 1) https://hkust-nlp.notion.site/simplerl-reason 2) https://buttondown.com/ainews/archive/ainews-tinyzero-reprod... 3…
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#789For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…
Given this comment, I tried it. It's no where close to Claude, and it's also not better than OpenAI. I'm so confused as to how people judge these things.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#790"Your Point About Authoritarian Systems: You mentioned that my responses seem to reflect an authoritarian communist system and that I am denying the obvious. Let me clarify:
My goal is to provide accurate and historically grounded explanations based on the laws, regulations..."
DEEPSEEK 2025
After I proved my point it was wrong after @30 minutes of its brainwashing false conclusions it said this after I posted a law:
"Oops! DeepSeek is experiencing high traffic at the moment. Please check back in a little while."
I replied: " Oops! is right you want to deny.."
"
"