Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

861–870 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#862

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

> For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini.

Worse at writing. Its prose is overwrought. It's yet to learn that "less is more"

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#863

Earlier quoted context omitted.

It’s not better than o1. And given that OpenAI is on the verge of releasing o3, has some “o4” in the pipeline, and Deepseek could only build this because of o1, I don’t think there’s as much competition as people seem to imply. I’m excited to see models become open, but given the curve of progress we’ve seen, even being “a little” behind is a gap that grows exponentially every day.

> It’s not better than o1. I thought that too before I used it to do real work.

Yes. It shines with real problems.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#864

Deepseek seems to create enormously long reasoning traces. I gave it the following for fun. It thought for a very long time (307 seconds), displaying a very long and stuttering trace before, losing confidence on the second part of the problem and getting it way wrong. GPTo1 got similarly tied in knots and took 193 seconds, getting the right order of magnitude for part 2 (0.001 inches). Gemini 2.0 Exp was much faster…

OpenAI reasoning traces are actually summarized by another model. The reason is that you can (as we are seeing happening now) “distill” the larger model reasoning into smaller models. Had OpenAI shown full traces in o1 answers they would have been giving gold to competition.

That's not the point of my post, but point taken.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#865

Commoditize your complement has been invoked as an explanation for Meta's strategy to open source LLM models (with some definition of "open" and "model"). Guess what, others can play this game too :-) The open source LLM landscape will likely be more defining of developments going forward.

Complement to which of Meta's products?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#866

Earlier quoted context omitted.

They're using it via fireworks.ai, which is the 685B model. https://fireworks.ai/models/fireworks/deepseek-r1

How do you know which version it is? I didn't see anything in that link.

An additional information panel shows up on the right hand side when you're logged in.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#867
Can someone share a youtube showing DeepSeek vs others? I glanced through comments and seeing lots of opinions, but no (easy) evidence. I would like to see a level of thoroughness that I could not do myself. Not naysaying one model over another, just good ole fashion elbow grease and scientific method for the layperson. I appreciate the help.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#868
post #862

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

> For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. Worse at writing. Its prose is overwrought. It's yet to learn that "less is more"

That's not what I've seen. See https://eqbench.com/results/creative-writing-v2/deepseek-ai_... , where someone fed it a large number of prompts.

Weirdly, while the first paragraph from the first story was barely GPT-3 grade, 99% of the rest of the output blew me away (and is continuing to do so, as I haven't finished reading it yet.)

I tried feeding a couple of the prompts to gpt-4o, o1-pro and the current Gemini 2.0 model, and the resulting output was nowhere near as well-crafted.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#869

Can someone share a youtube showing DeepSeek vs others? I glanced through comments and seeing lots of opinions, but no (easy) evidence. I would like to see a level of thoroughness that I could not do myself. Not naysaying one model over another, just good ole fashion elbow grease and scientific method for the layperson. I appreciate the help.

Here [1] is the leaderboard from chabot arena, where users vote on the output of two anonymous models. Deepseek R1 needs more data points- but it already climbed to No 1 with Style control ranking, which is pretty impressive.

Link [2] to the result on more standard LLM benchmarks. They conveniently placed the results on the first page of the paper.

[1] https://lmarena.ai/?leaderboard

[2] https://arxiv.org/pdf/2501.12948 (PDF)

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#870
post #665

Deepseek R1 now has almost 1M downloads in Ollama: https://ollama.com/library/deepseek-r1 That is a lot of people running their own models. OpenAI is probably is panic mode right now.

most of those models aren’t r1

they are distillations of r1, and work fairly well given the modest hardware they need.
Post reply on HN