Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

401–410 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#401
post #357

DeepSeek V3 came in the perfect time, precisely when Claude Sonnet turned into crap and barely allows me to complete something without me hitting some unexpected constraints. Idk, what their plans is and if their strategy is to undercut the competitors but for me, this is a huge benefit. I received 10$ free credits and have been using Deepseeks api a lot, yet, I have barely burned a single dollar, their pricing are t…

Prices will increase by five times in February, but it will still be extremely cheap compared to Sonnet. $15/million vs $1.10/million for output is a world of difference. There is no reason to stop using Sonnet, but I will probably only use it when DeepSeek goes into a tailspin or I need extra confidence in the responses.

Could this trend bankrupt most incumbent LLM companies?

They’ve invested billions on their models and infrastructure, which they need to recover through revenue

If new exponentially cheaper models/services come out fast enough, the incumbent might not be able to recover their investments

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#402

Earlier quoted context omitted.

Great as long as you’re not interested in Tiananmen Square or the Uighurs.

try asking US models about the influence of Israeli diaspora on funding genocide in Gaza then come back

[dead]

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#403
post #125

"Reasoning" will be disproven for this again within a few days I guess. Context: o1 does not reason, it pattern matches. If you rename variables, suddenly it fails to solve the request.

Rename to equally reasonable variable names, or to intentionally misleading or meaningless ones? Good naming is one of the best ways to make reading unfamiliar code easier for people, don't see why actual AGI wouldn't also get tripped up there.

Can't we sometimed expect more from computers than people, especially around something that compilers have done for decades.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#404

Earlier quoted context omitted.

Outside of Veo2 - which I can’t access anyway - they’re definitely ahead in AI video gen

the big american labs don’t care about ai video gen

They didn't care about neural networks once.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#405

Earlier quoted context omitted.

There's an interesting tweet here from someone who used to work at DeepSeek, which describes their hiring process and culture. No mention of LeetCoding for sure! https://x.com/wzihanw/status/1872826641518395587

they almost certainly ask coding/technical questions. the people doing this work are far beyond being gatekept by leetcode leetcode is like HN’s “DEI” - something they want to blame everything on

Did you read the tweet? It doesn't sound that way to me. They hire specialized talent (note especially the "Know-It-All" part)

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#406

I've always been leery about outrageous GPU investments, at some point I'll dig through and find my prior comments where I've said as much to that effect. The CEOs, upper management, and governments derive their importance on how much money they can spend - AI gave them the opportunity for them to confidently say that if you give me $X I can deliver Y and they turn around and give that money to NVidia. The problem wa…

The cost of having excess compute is less than the cost of not having enough compute to be competitive. Because of demand, if you realize you your current compute is insufficient there is a long turnaround to building up your infrastructure, at which point you are falling behind. All the major players are simultaneously working on increasing capabilities and reducing inference cost. What they aren’t optimizing is the…

As long as you have investors shovelling money in.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#407

Curious if this will prompt OpenAI to unveil o1’s “thinking” steps. Afaict they’ve hidden them primarily to stifle the competition… which doesn’t seem to matter at present!

The thinking steps for o1 have been recently improved.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#408
post #203

Earlier quoted context omitted.

Sigh, I don't understand why they had to do the $500 billion announcement with the president. So many people now wrongly think Trump just gave OpenAI $500 billion of the taxpayers' money.

It means he’ll knock down regulatory barriers and mess with competitors because his brand is associated with it. It was a smart poltical move by OpenAI.

Until the regime is toppled, then it will look very short-sighted and stupid.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#409
post #16

Earlier quoted context omitted.

I've been confused over this. I've seen a $5.5M # for training, and commensurate commentary along the lines of what you said, but it elides the cost of the base model AFAICT.

With $5.5M, you can buy around 150 H100s. Experts correct me if I’m wrong but it’s practically impossible to train a model like that with that measly amount. So I doubt that figure includes all the cost of training.

Is it a fine tune effectively?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#410

Even if you think this particular team cheated, the idea that nobody will find ways of making training more efficient seems silly - these huge datacenter investments for purely AI will IMHO seem very short sighted in 10 years

More like three years. Even in the best case the retained value curve of GPUs is absolutely terrible. Most of these huge investments in GPUs are going to be massive losses.

GPUs can do other stuff though. I wouldn't bet on GPU ghost towns.
Post reply on HN