Earlier quoted context omitted.
Are you referring to the distilled models?
yes, they are not r1
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
861–870 of 1001 posts
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#862For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…
Worse at writing. Its prose is overwrought. It's yet to learn that "less is more"
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#863Earlier quoted context omitted.
It’s not better than o1. And given that OpenAI is on the verge of releasing o3, has some “o4” in the pipeline, and Deepseek could only build this because of o1, I don’t think there’s as much competition as people seem to imply. I’m excited to see models become open, but given the curve of progress we’ve seen, even being “a little” behind is a gap that grows exponentially every day.
> It’s not better than o1. I thought that too before I used it to do real work.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#864Deepseek seems to create enormously long reasoning traces. I gave it the following for fun. It thought for a very long time (307 seconds), displaying a very long and stuttering trace before, losing confidence on the second part of the problem and getting it way wrong. GPTo1 got similarly tied in knots and took 193 seconds, getting the right order of magnitude for part 2 (0.001 inches). Gemini 2.0 Exp was much faster…
OpenAI reasoning traces are actually summarized by another model. The reason is that you can (as we are seeing happening now) “distill” the larger model reasoning into smaller models. Had OpenAI shown full traces in o1 answers they would have been giving gold to competition.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#865Commoditize your complement has been invoked as an explanation for Meta's strategy to open source LLM models (with some definition of "open" and "model"). Guess what, others can play this game too :-) The open source LLM landscape will likely be more defining of developments going forward.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#866Earlier quoted context omitted.
They're using it via fireworks.ai, which is the 685B model. https://fireworks.ai/models/fireworks/deepseek-r1
How do you know which version it is? I didn't see anything in that link.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#867Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#868For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…
> For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. Worse at writing. Its prose is overwrought. It's yet to learn that "less is more"
Weirdly, while the first paragraph from the first story was barely GPT-3 grade, 99% of the rest of the output blew me away (and is continuing to do so, as I haven't finished reading it yet.)
I tried feeding a couple of the prompts to gpt-4o, o1-pro and the current Gemini 2.0 model, and the resulting output was nowhere near as well-crafted.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#869Can someone share a youtube showing DeepSeek vs others? I glanced through comments and seeing lots of opinions, but no (easy) evidence. I would like to see a level of thoroughness that I could not do myself. Not naysaying one model over another, just good ole fashion elbow grease and scientific method for the layperson. I appreciate the help.
Link [2] to the result on more standard LLM benchmarks. They conveniently placed the results on the first page of the paper.
[1] https://lmarena.ai/?leaderboard
[2] https://arxiv.org/pdf/2501.12948 (PDF)
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#870Deepseek R1 now has almost 1M downloads in Ollama: https://ollama.com/library/deepseek-r1 That is a lot of people running their own models. OpenAI is probably is panic mode right now.
most of those models aren’t r1