Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

491–500 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#491
post #268

I've been using https://chat.deepseek.com/ over My ChatGPT Pro subscription because being able to read the thinking in the way they present it is just much much easier to "debug" - also I can see when it's bending it's reply to something, often softening it or pandering to me - I can just say "I saw in your thinking you should give this type of reply, don't do that". If it stays free and gets better that's going to b…

The chain of thought is super useful in so many ways, helping me: (1) learn, way beyond the final answer itself, (2) refine my prompt, whether factually or stylistically, (3) understand or determine my confidence in the answer.

do you have any resources related to these???

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#492
post #269
post #262

Earlier quoted context omitted.

> i.e. high speed rail network instead You want to invest $500B to a high speed rail network which the Chinese could build for $50B?

Just commission the Chinese and make it 10X bigger then. In the case of the AI, they appear to commission Sam Altman and Larry Ellison.

It doesn't matter who you "commission" to do the actual work, most of the additional cost is in legal battles over rights of way and environmental impacts and other things that are independent of the construction work.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#493
post #79

How can openai justify their $200/mo subscriptions if a model like this exists at an incredibly low price point? Operator? I've been impressed in my brief personal testing and the model ranks very highly across most benchmarks (when controlled for style it's tied number one on lmarena). It's also hilarious that openai explicitly prevented users from seeing the CoT tokens on the o1 model (which you still pay for btw)…

From my casual read, right now everyone is on reputation tarnishing tirade, like spamming “Chinese stealing data! Definitely lying about everything! API can’t be this cheap!”. If that doesn’t go through well, I’m assuming lobbyism will start for import controls, which is very stupid. I have no idea how they can recover from it, if DeepSeek’s product is what they’re advertising.

Funny, everything I see (not actively looking for DeepSeek related content) is absolutely raving about it and talking about it destroying OpenAI (random YouTube thumbnails, most comments in this thread, even CNBC headlines).

If DeepSeek's claims are accurate, then they themselves will be obsolete within a year, because the cost to develop models like this has dropped dramatically. There are going to be a lot of teams with a lot of hardware resources with a lot of motivation to reproduce and iterate from here.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#494
post #347

Earlier quoted context omitted.

I never said Llama is mediocre. I said the teams they put together is full of people chasing money. And the billions Meta is burning is going straight to mediocrity. They’re bloated. And we know exactly why Meta is doing this and it’s not because they have some grand scheme to build up AI. It’s to keep these people away from their competition. Same with billions in GPU spend. They want to suck up resources away from…

> And we know exactly why Meta is doing this and it’s not because they have some grand scheme to build up AI. It’s to keep these people away from their competition I don't see how you can confidently say this when AI researchers and engineers are remunerated very well across the board and people are moving across companies all the time, if the plan is as you described it, it is clearly not working. Zuckerberg seems c…

this is the same magical thinking Uber had when they were gonna have self driving cars replace their drivers

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#495
post #430
post #401

Earlier quoted context omitted.

Could this trend bankrupt most incumbent LLM companies? They’ve invested billions on their models and infrastructure, which they need to recover through revenue If new exponentially cheaper models/services come out fast enough, the incumbent might not be able to recover their investments

I literally cannot see how OpenAI and Anthropic can justify their valuation given DeepSeek. In business, if you can provide twice the value at half the price, you will destroy the incumbent. Right now, DeepSeek is destroying on price and provides somewhat equivalent value compared to Sonnet. I still believe Sonnet is better, but I don't think it is 10 times better. Something else that DeepSeek can do, which I am not…

> I still believe Sonnet is better, but I don't think it is 10 times better.

Sonnet doesn't need to be 10 times better. It just needs to be better enough such that the downstream task improves more than the additional cost.

This is a much more reasonable hurdle. If you're able to improve the downstream performance of something that costs $500k/year by 1% then the additional cost of Sonnet just has to be less than $5k/year for there to be positive ROI.

I'm a big fan of DeepSeek. And the VC funded frontier labs may be screwed. But I don't think R1 is terminal for them. It's still a very competitive field.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#496
post #182

Earlier quoted context omitted.

Correct me if I'm wrong but if Chinese can produce the same quality at %99 discount, then the supposed $500B investment is actually worth $5B. Isn't that the kind wrong investment that can break nations? Edit: Just to clarify, I don't imply that this is public money to be spent. It will commission $500B worth of human and material resources for 5 years that can be much more productive if used for something else - i.e…

500 billion can move whole country to renewable energy

Not even close. The US spends roughly $2trillion/year on energy. If you assume 10% return on solar, that's $20trillion of solar to move the country to renewable. That doesn't calculate the cost of batteries which probably will be another $20trillion.

Edit: asked Deepseek about it. I was kinda spot on =)

Cost Breakdown

Solar Panels $13.4–20.1 trillion (13,400 GW × $1–1.5M/GW)

Battery Storage $16–24 trillion (80 TWh × $200–300/kWh)

Grid/Transmission $1–2 trillion

Land, Installation, Misc. $1–3 trillion

Total $30–50 trillion

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#497
post #205

Earlier quoted context omitted.

I don't say that at all. Money spent on BS still sucks resources, no matter who spends that money. They are not going to make the GPU's from 500 billion dollar banknotes, they will pay people $500B to work on this stuff which means people won't be working on other stuff that can actually produce value worth more than the $500B. I guess the power plants are salvageable.

By that logic all money is waste. The money isnt destroyed when it is spent. It is transferred into someone else's bank account only. This process repeats recursively until taxation returns all money back to the treasury to be spent again. And out of this process of money shuffling: entire nations full of power plants!

Money can be destroyed with inflation.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#498
post #115

Earlier quoted context omitted.

have you tried asking chatgpt something even slightly controversial? chatgpt censors much more than deepseek does. also deepseek is open-weights. there is nothing preventing you from doing a finetune that removes the censorship. they did that with llama2 back in the day.

> chatgpt censors much more than deepseek does This is an outrageous claim with no evidence, as if there was any equivalence between government enforced propaganda and anything else. Look at the system prompts for DeepSeek and it’s even more clear. Also: fine tuning is not relevant when what is deployed at scale brainwashes the masses through false and misleading responses.

why do you lie, it is blatantly obvious chatgpt censors a ton of things and has a bit of left-tilt too while trying hard to stay neutral.

If you think these tech companies are censoring all of this “just because” and instead of being completely torched by the media, and government who’ll use it as an excuse to take control of AI, then you’re sadly lying to yourself.

Think about it for a moment, why did Trump (and im not a trump supporter) re-appeal Biden’s AI Executive Order 2023 ? , what was in it ? , it is literally a propaganda enforcement article, written in sweet sounding, well meaning words.

It’s ok, no country is angel, even the american founding fathers would except americans to be critical of its government during moments, there’s no need for thinking that America = Good and China = Bad. We do have a ton of censorship in the “free world” too and it is government enforced, or else you wouldnt have seen so many platforms turn the tables on moderation, the moment trump got elected, the blessing for censorship directly comes from government.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#499

Earlier quoted context omitted.

It does if the spend drives GPU prices so high that more researchers can't afford to use them. And DS demonstrated what a small team of researchers can do with a moderate amount of GPUs.

The DS team themselves suggest large amounts of compute are still required

However, look at the figure for R1-zero. The x-axis is effectively the number of RL steps, measured in the thousands. Each of them involves a whole group of inferences, but compare that to the gradient updates required for consuming 15 trillion tokens during pretraining, and it is still a bargain. Direct RL on the smaller models was not effective as quickly as with DeepSeek v3, so although in principle it might work at some level of compute, it was much cheaper to do SFT of these small models using reasoning traces of the big model. The distillation SFT on 800k example traces probably took much less than 0.1% of the pretraining compute of these smaller models, so this is the compute budget they compare RL against in the snippet that you quote.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#500
For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini.

It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc.

We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now.

The justification for keeping the sauce secret just seems a lot more absurd. None of the top secret sauce that those companies have been hyping up is worth anything now that there is a superior open source model. Let that sink in.

This is real competition. If we can't have it in EVs at least we can have it in AI models!

Post reply on HN