For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…
Here Deepseek r1 fixes a python bug. Its fix is the same as the original code. I have not seen that level of stupidity from o1 or sonnet 3.5 https://x.com/alecm3/status/1883147247485170072?t=55xwg97roj...
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
641–650 of 1001 posts
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#642Earlier quoted context omitted.
False equivalency. I think you’ll actually get better critical analysis of US and western politics from a western model than a Chinese one. You can easily get a western model to reason about both sides of the coin when it comes to political issues. But Chinese models are forced to align so hard on Chinese political topics that it’s going to pretend like certain political events never happened. E.g try getting them to…
This is not really my experience with western models. I am not from the US though, so maybe what you consider a balanced perspective or reasoning about both sides is not the same as what I would call one. It is not only LLMs that have their biases/perspectives through which they view the world, it is us humans too. The main difference imo is not between western and chinese models but between closed and, in whichever…
I'm also not from the US, but I'm not sure what you mean here. Unless you're talking about defaulting to answer in Imperial units, or always using examples from the US, which is a problem the entire English speaking web has.
Can you give some specific examples of prompts that will demonstrate the kind of Western bias or censorship you're talking about?
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#643"Prove or disprove: there exists a closed, countable, non-trivial partition of a connected Hausdorff space."
And it made a pretty amateurish mistake:
"Thus, the real line R with the partition {[n,n+1]∣n∈Z} serves as a valid example of a connected Hausdorff space with a closed, countable, non-trivial partition."
o1 gets this prompt right the few times I tested it (disproving it using something like Sierpinski).
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#644Earlier quoted context omitted.
I tried Deepseek R1 via Kagi assistant and it was much better than claude or gpt. I asked for suggestions for rust libraries for a certain task and the suggestions from Deepseek were better. Results here: https://x.com/larrysalibra/status/1883016984021090796
This is really poor test though, of course the most recently trained model knows the newest libraries or knows that a library was renamed. Not disputing it's best at reasoning but you need a different test for that.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#645we've been tracking the deepseek threads extensively in LS. related reads: - i consider the deepseek v3 paper required preread https://github.com/deepseek-ai/DeepSeek-V3 - R1 + Sonnet > R1 or O1 or R1+R1 or O1+Sonnet or any other combo https://aider.chat/2025/01/24/r1-sonnet.html - independent repros: 1) https://hkust-nlp.notion.site/simplerl-reason 2) https://buttondown.com/ainews/archive/ainews-tinyzero-reprod... 3…
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#646Earlier quoted context omitted.
I am not surprised if US Govt would mandate "Tiananmen-test" for LLMs in the future to have "clean LLM". Anyone working for federal govt or receiving federal money would only be allowed to use "clean LLM"
Curious to learn what do you think would be a good "Tiananmen-test" for US based models
That's it
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#647Earlier quoted context omitted.
Oh, my experience was different. Got the model through ollama. I'm quite impressed how they managed to bake in the censorship. It's actually quite open about it. I guess censorship doesnt have as bad a rep in china as it has here? So it seems to me that's one of the main achievements of this model. Also another finger to anyone who said they can't publish their models cause of ethical reasons. Deepseek demonstrated c…
I mean US models are highly censored too.
False equivalency if you ask me. There may be some alignment to make the models polite and avoid outright racist replies and such. But political censorship? Please elaborate
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#648Genuinly curious, what is everyone using reasoning models for? (R1/o1/o3)
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#649Earlier quoted context omitted.
500 billion can move whole country to renewable energy
Not even close. The US spends roughly $2trillion/year on energy. If you assume 10% return on solar, that's $20trillion of solar to move the country to renewable. That doesn't calculate the cost of batteries which probably will be another $20trillion. Edit: asked Deepseek about it. I was kinda spot on =) Cost Breakdown Solar Panels $13.4–20.1 trillion (13,400 GW × $1–1.5M/GW) Battery Storage $16–24 trillion (80 TWh ×…
The most common idea is to spend 3-5% of GDP per year for the transition (750-1250 bn USD per year for the US) over the next 30 years. Certainly a significant sum, but also not too much to shoulder.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#650Earlier quoted context omitted.
> Deepseek could only build this because of o1, I don’t think there’s as much competition as people seem to imply And this is based on what exactly? OpenAI hides the reasoning steps, so training a model on o1 is very likely much more expensive (and much less useful) than just training it directly on a cheaper model.
Because literally before o1, no one is doing COT style test time scaling. It is a new paradigm. The talking point back then, is the LLM hits the wall. R1's biggest contribution IMO, is R1-Zero, I am fully sold with this they don't need o1's output to be as good. But yeah, o1 is still the herald.
Like, this idea always seemed completely obvious to me, and I figured the only reason why it hadn't been done yet is just because (at the time) models weren't good enough. (So it just caused them to get confused, and it didn't improve results.)
Presumably OpenAI were the first to claim this achievement because they had (at the time) the strongest model (+ enough compute). That doesn't mean COT was a revolutionary idea, because imo it really wasn't. (Again, it was just a matter of having a strong enough model, enough context, enough compute for it to actually work. That's not an academic achievement, just a scaling victory.)