Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

221–230 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#221
post #182

DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...

Correct me if I'm wrong but if Chinese can produce the same quality at %99 discount, then the supposed $500B investment is actually worth $5B. Isn't that the kind wrong investment that can break nations? Edit: Just to clarify, I don't imply that this is public money to be spent. It will commission $500B worth of human and material resources for 5 years that can be much more productive if used for something else - i.e…

if you say, i wanna build 5 nuclear reactors and I need 200 billion $$. I would believe it because, you can ballpark it with some stats.

For tech like LLMs, it feels irresponsible to say 500 billion $$ investment and then place that into R&D. What if in 2026, we realize we can create it for 2 billion$, and let the 498 billion $ sitting in a few consumers.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#222
post #205
post #203

Earlier quoted context omitted.

Sigh, I don't understand why they had to do the $500 billion announcement with the president. So many people now wrongly think Trump just gave OpenAI $500 billion of the taxpayers' money.

I don't say that at all. Money spent on BS still sucks resources, no matter who spends that money. They are not going to make the GPU's from 500 billion dollar banknotes, they will pay people $500B to work on this stuff which means people won't be working on other stuff that can actually produce value worth more than the $500B. I guess the power plants are salvageable.

By that logic all money is waste. The money isnt destroyed when it is spent. It is transferred into someone else's bank account only. This process repeats recursively until taxation returns all money back to the treasury to be spent again. And out of this process of money shuffling: entire nations full of power plants!

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#223
post #218

I was completely surprised that the reasoning comes from within the model. When using gpt-o1 I thought it's actually some optimized multi-prompt chain, hidden behind an API endpoint. Something like: collect some thoughts about this input; review the thoughts you created; create more thoughts if needed or provide a final answer; ...

I think the reason why it works is also because chain-of-thought (CoT), in the original paper by Denny Zhou et. al, worked from "within". The observation was that if you do CoT, answers get better.

Later on community did SFT on such chain of thoughts. Arguably, R1 shows that was a side distraction, and instead a clean RL reward would've been better suited.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#224

Earlier quoted context omitted.

I guess all that leetcoding and stack ranking didn't in fact produce "the cream of the crop"...

It produces the cream of the leetcoding stack ranking crop.

You get what you measure.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#225
post #186

Earlier quoted context omitted.

At least it’s not home grown propaganda from the US, so will likely not cover most other topics of interest.

What are you basing this whataboutism on?

Not a fan of censorship here, but Chinese models are (subjectively) less propagandized than US models. If you ask US models about China, for instance, they'll tend towards the antagonistic perspective favored by US media. Chinese models typically seem to take a more moderate, considered tone when discussing similar subjects. US models also suffer from safety-based censorship, especially blatant when "safety" involves protection of corporate resources (eg. not helping the user to download YouTube videos).

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#226
post #219

Earlier quoted context omitted.

$500 billion is $500 billion. If new technology means we can get more for a dollar spent, then $500 billion gets more, not less.

That's right but the money is given to the people who do it for $500B and there are much better ones who can do it for $5B instead and if they end up getting $6B they will have a better model. What now?

I don't know how to answer this because these are arbitrary numbers.

The money is not spent. Deepseek published their methodology, incumbents can pivot and build on it. No one knows what the optimal path is, but we know it will cost more.

I can assure you that OpenAI won't continue to produce inferior models at 100x the cost.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#227
post #79

How can openai justify their $200/mo subscriptions if a model like this exists at an incredibly low price point? Operator? I've been impressed in my brief personal testing and the model ranks very highly across most benchmarks (when controlled for style it's tied number one on lmarena). It's also hilarious that openai explicitly prevented users from seeing the CoT tokens on the o1 model (which you still pay for btw)…

I find that this model feels more human, purely because of the reasoning style (first person). In its reasoning text, it comes across as a neurotic, eager to please smart “person”, which is hard not to anthropomorphise

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#228
post #182

Earlier quoted context omitted.

Correct me if I'm wrong but if Chinese can produce the same quality at %99 discount, then the supposed $500B investment is actually worth $5B. Isn't that the kind wrong investment that can break nations? Edit: Just to clarify, I don't imply that this is public money to be spent. It will commission $500B worth of human and material resources for 5 years that can be much more productive if used for something else - i.e…

if you say, i wanna build 5 nuclear reactors and I need 200 billion $$. I would believe it because, you can ballpark it with some stats. For tech like LLMs, it feels irresponsible to say 500 billion $$ investment and then place that into R&D. What if in 2026, we realize we can create it for 2 billion$, and let the 498 billion $ sitting in a few consumers.

Don’t think of it as “spend a fixed amount to get a fixed outcome”. Think of it as “spend a fixed amount and see how far you can get”

It may still be flawed or misguided or whatever, but it’s not THAT bad.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#229

...and China is two years behind in AI. Right ?

They were 6 months behind US frontier until deepseek r1. Now maybe 4? It's hard to say.

Outside of Veo2 - which I can’t access anyway - they’re definitely ahead in AI video gen

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#230
post #219

Earlier quoted context omitted.

$500 billion is $500 billion. If new technology means we can get more for a dollar spent, then $500 billion gets more, not less.

That's right but the money is given to the people who do it for $500B and there are much better ones who can do it for $5B instead and if they end up getting $6B they will have a better model. What now?

Are you under the impression it was some kind of fixed-scope contractor bid for a fixed price?
Post reply on HN