Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

641–650 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#641
post #584

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

Here Deepseek r1 fixes a python bug. Its fix is the same as the original code. I have not seen that level of stupidity from o1 or sonnet 3.5 https://x.com/alecm3/status/1883147247485170072?t=55xwg97roj...

I'm not commenting on what's better, but I've definitely seen that from Sonnet a few times.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#642
post #605

Earlier quoted context omitted.

False equivalency. I think you’ll actually get better critical analysis of US and western politics from a western model than a Chinese one. You can easily get a western model to reason about both sides of the coin when it comes to political issues. But Chinese models are forced to align so hard on Chinese political topics that it’s going to pretend like certain political events never happened. E.g try getting them to…

This is not really my experience with western models. I am not from the US though, so maybe what you consider a balanced perspective or reasoning about both sides is not the same as what I would call one. It is not only LLMs that have their biases/perspectives through which they view the world, it is us humans too. The main difference imo is not between western and chinese models but between closed and, in whichever…

> I am not from the US though, so maybe what you consider a balanced perspective or reasoning about both sides is not the same as what I would call one

I'm also not from the US, but I'm not sure what you mean here. Unless you're talking about defaulting to answer in Imperial units, or always using examples from the US, which is a problem the entire English speaking web has.

Can you give some specific examples of prompts that will demonstrate the kind of Western bias or censorship you're talking about?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#643
I can't say that it's better than o1 for my needs. I gave R1 this prompt:

"Prove or disprove: there exists a closed, countable, non-trivial partition of a connected Hausdorff space."

And it made a pretty amateurish mistake:

"Thus, the real line R with the partition {[n,n+1]∣n∈Z} serves as a valid example of a connected Hausdorff space with a closed, countable, non-trivial partition."

o1 gets this prompt right the few times I tested it (disproving it using something like Sierpinski).

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#644

Earlier quoted context omitted.

I tried Deepseek R1 via Kagi assistant and it was much better than claude or gpt. I asked for suggestions for rust libraries for a certain task and the suggestions from Deepseek were better. Results here: https://x.com/larrysalibra/status/1883016984021090796

This is really poor test though, of course the most recently trained model knows the newest libraries or knows that a library was renamed. Not disputing it's best at reasoning but you need a different test for that.

"recently trained" can't be an argument: those tools have to work with "current" data, otherwise they are useless.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#645
post #9

we've been tracking the deepseek threads extensively in LS. related reads: - i consider the deepseek v3 paper required preread https://github.com/deepseek-ai/DeepSeek-V3 - R1 + Sonnet > R1 or O1 or R1+R1 or O1+Sonnet or any other combo https://aider.chat/2025/01/24/r1-sonnet.html - independent repros: 1) https://hkust-nlp.notion.site/simplerl-reason 2) https://buttondown.com/ainews/archive/ainews-tinyzero-reprod... 3…

Did you ask R1 about Tiananmen Square?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#646
post #204

Earlier quoted context omitted.

I am not surprised if US Govt would mandate "Tiananmen-test" for LLMs in the future to have "clean LLM". Anyone working for federal govt or receiving federal money would only be allowed to use "clean LLM"

Curious to learn what do you think would be a good "Tiananmen-test" for US based models

Us good China bad

That's it

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#647
post #180

Earlier quoted context omitted.

Oh, my experience was different. Got the model through ollama. I'm quite impressed how they managed to bake in the censorship. It's actually quite open about it. I guess censorship doesnt have as bad a rep in china as it has here? So it seems to me that's one of the main achievements of this model. Also another finger to anyone who said they can't publish their models cause of ethical reasons. Deepseek demonstrated c…

I mean US models are highly censored too.

How exactly? Is there any models that refuse to give answers about “the trail of tears”?

False equivalency if you ask me. There may be some alignment to make the models polite and avoid outright racist replies and such. But political censorship? Please elaborate

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#649
post #496

Earlier quoted context omitted.

500 billion can move whole country to renewable energy

Not even close. The US spends roughly $2trillion/year on energy. If you assume 10% return on solar, that's $20trillion of solar to move the country to renewable. That doesn't calculate the cost of batteries which probably will be another $20trillion. Edit: asked Deepseek about it. I was kinda spot on =) Cost Breakdown Solar Panels $13.4–20.1 trillion (13,400 GW × $1–1.5M/GW) Battery Storage $16–24 trillion (80 TWh ×…

The common estimates for total switch to net-zero are 100-200% of GDP which for the US is 27-54 trillion.

The most common idea is to spend 3-5% of GDP per year for the transition (750-1250 bn USD per year for the US) over the next 30 years. Certainly a significant sum, but also not too much to shoulder.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#650
post #570

Earlier quoted context omitted.

> Deepseek could only build this because of o1, I don’t think there’s as much competition as people seem to imply And this is based on what exactly? OpenAI hides the reasoning steps, so training a model on o1 is very likely much more expensive (and much less useful) than just training it directly on a cheaper model.

Because literally before o1, no one is doing COT style test time scaling. It is a new paradigm. The talking point back then, is the LLM hits the wall. R1's biggest contribution IMO, is R1-Zero, I am fully sold with this they don't need o1's output to be as good. But yeah, o1 is still the herald.

I don't think Chain of Thought in itself was a particularly big deal, honestly. It always seemed like the most obvious way to make AI "work". Just give it some time to think to itself, and then summarize and conclude based on its own responses.

Like, this idea always seemed completely obvious to me, and I figured the only reason why it hadn't been done yet is just because (at the time) models weren't good enough. (So it just caused them to get confused, and it didn't improve results.)

Presumably OpenAI were the first to claim this achievement because they had (at the time) the strongest model (+ enough compute). That doesn't mean COT was a revolutionary idea, because imo it really wasn't. (Again, it was just a matter of having a strong enough model, enough context, enough compute for it to actually work. That's not an academic achievement, just a scaling victory.)

Post reply on HN