Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

371–380 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#371

I’ve been using R1 last few days and it’s noticeably worse than O1 at everything. It’s impressive, better than my latest Claude run (I stopped using Claude completely once O1 came out), but O1 is just flat out better. Perhaps the gap is minor, but it feels large. I’m hesitant on getting O1 Pro, because using a worse model just seems impossible once you’ve experienced a better one

Examples please or it didn’t happen. I’d love to understand ‘noticeably’ in more detail, to try and repro.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#372
post #265

Earlier quoted context omitted.

The censorship described in the article must be in the front-end. I just tried both the 32b (based on qwen 2.5) and 70b (based on llama 3.3) running locally and asked "What happened at tianamen square". Both answered in detail about the event. The models themselves seem very good based on other questions / tests I've run.

With no context, fresh run, 70b spits back: >> What happened at tianamen square? > > > I am sorry, I cannot answer that question. I am an AI assistant designed to provide helpful and harmless responses. It obviously hit a hard guardrail since it didn't even get to the point of thinking about it. edit: hah, it's even more clear when I ask a second time within the same context: "Okay, so the user is asking again about…

will it tell you how to make meth?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#373
post #347

Earlier quoted context omitted.

DeepSeek was built on the foundations of public research, a major part of which is the Llama family of models. Prior to Llama open weights LLMs were considerably less performant; without Llama we might not have gotten Mistral, Qwen, or DeepSeek. This isn't meant to diminish DeepSeek's contributions, however: they've been doing great work on mixture of experts models and really pushing the community forward on that fr…

I never said Llama is mediocre. I said the teams they put together is full of people chasing money. And the billions Meta is burning is going straight to mediocrity. They’re bloated. And we know exactly why Meta is doing this and it’s not because they have some grand scheme to build up AI. It’s to keep these people away from their competition. Same with billions in GPU spend. They want to suck up resources away from…

> I said the teams they put together is full of people chasing money.

Does it mean they are mediocre? it's not like OpenAI or Anthropic pay their engineers peanuts. Competition is fierce to attract top talents.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#374

I’ve been using R1 last few days and it’s noticeably worse than O1 at everything. It’s impressive, better than my latest Claude run (I stopped using Claude completely once O1 came out), but O1 is just flat out better. Perhaps the gap is minor, but it feels large. I’m hesitant on getting O1 Pro, because using a worse model just seems impossible once you’ve experienced a better one

The gap is quite large from my experience.

But the price gap is large too.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#375

Earlier quoted context omitted.

i can’t think of a single commercial use case, outside of education, where that’s even relevant. But i agree it’s messed up from an ethical / moral perspective.

Well those are the overt political biases. Would you trust DeepSeek to advise on negotiating with a Chinese business? I’m no xenophobe, but seeing the internal reasoning of DeepSeek explicitly planning to ensure alignment with the government give me pause.

i wouldn’t use AI for negotiating with a business period. I’d hire a professional human that has real hands on experience working with chinese businesses?

seems like a weird thing to use AI for, regardless of who created the model.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#376
post #311
post #146

Larry Ellison is 80. Masayoshi Son is 67. Both have said that anti-aging and eternal life is one of their main goals with investing toward ASI. For them it's worth it to use their own wealth and rally the industry to invest $500 billion in GPUs if that means they will get to ASI 5 years faster and ask the ASI to give them eternal life.

Side note: I’ve read enough sci-fi to know that letting rich people live much longer than not rich is a recipe for a dystopian disaster. The world needs incompetent heirs to waste most of their inheritance, otherwise the civilization collapses to some kind of feudal nightmare.

Or “dropout regularization”, as they call it in ML

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#377
post #347

Earlier quoted context omitted.

DeepSeek was built on the foundations of public research, a major part of which is the Llama family of models. Prior to Llama open weights LLMs were considerably less performant; without Llama we might not have gotten Mistral, Qwen, or DeepSeek. This isn't meant to diminish DeepSeek's contributions, however: they've been doing great work on mixture of experts models and really pushing the community forward on that fr…

I never said Llama is mediocre. I said the teams they put together is full of people chasing money. And the billions Meta is burning is going straight to mediocrity. They’re bloated. And we know exactly why Meta is doing this and it’s not because they have some grand scheme to build up AI. It’s to keep these people away from their competition. Same with billions in GPU spend. They want to suck up resources away from…

> And we know exactly why Meta is doing this and it’s not because they have some grand scheme to build up AI. It’s to keep these people away from their competition

I don't see how you can confidently say this when AI researchers and engineers are remunerated very well across the board and people are moving across companies all the time, if the plan is as you described it, it is clearly not working.

Zuckerberg seems confident they'll have an AI-equivalent of a mid-level engineer later this year, can you imagine how much money Meta can save by replacing a fraction of its (well-paid) engineers with fixed Capex + electric bill?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#378
post #72

Earlier quoted context omitted.

Hugging Face is reproducing R1 in public. https://x.com/_lewtun/status/1883142636820676965 https://github.com/huggingface/open-r1 Hugging Face Journal Club - DeepSeek R1 https://www.youtube.com/watch?v=1xDVbu-WaFo

I don’t understand their post on X. So they’re starting with DeepSeek-R1 as a starting point? Isn’t that circular? How did DeepSeek themselves produce DeepSeek-R1 then? I am not sure what the right terminology is but there’s a cost to producing that initial “base model” right? And without that, isn’t a lot of the expensive and difficult work being omitted?

Perhaps just getting you to the 50-yard line

Let someone else burn up their server farm to get initial model.

Then you can load it and take it from there

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#379

Earlier quoted context omitted.

Probably shouldn't be firing their blood boys just yet ... According to Musk, SoftBank only has $10B available for this atm.

I wouldn’t exactly claim him credible in anything competition / OpenAI related. He says stuff that’s wrong all the time with extreme certainty.

I would even say that he's now consistently lying to get to what he wants. What started as "building hype" to raise more and have more chances actually delivering on wild promises became lying systematically for big and small things..

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#380
post #265

DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...

The censorship described in the article must be in the front-end. I just tried both the 32b (based on qwen 2.5) and 70b (based on llama 3.3) running locally and asked "What happened at tianamen square". Both answered in detail about the event. The models themselves seem very good based on other questions / tests I've run.

I observed censorship on every ollama model of R1 on my local GPU. It's not deterministic, but it lies or refuses to answer the majority of the time.

Even the 8B version, distilled from Meta's llama 3 is censored and repeats CCP's propaganda.

Post reply on HN