Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

341–350 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#341
post #204
post #151

Aside from the usual Tiananmen Square censorship, there's also some other propaganda baked-in: https://prnt.sc/HaSc4XZ89skA (from reddit)

I am not surprised if US Govt would mandate "Tiananmen-test" for LLMs in the future to have "clean LLM". Anyone working for federal govt or receiving federal money would only be allowed to use "clean LLM"

Curious to learn what do you think would be a good "Tiananmen-test" for US based models

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#342

DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...

> As Nvidia senior research manager Jim Fan put it on X: “We are living in a timeline where a non-US company is keeping the original mission of OpenAI alive — truly open, frontier research that empowers all. . ."

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#343
post #59

Earlier quoted context omitted.

Regular coding questions mostly. For me o1 generally gives better code and understands the prompt more completely (haven’t started using r1 or o3 regularly enough to opine).

o3 isn’t available

Right, and that doesn't contradict what I wrote.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#344
post #340

Earlier quoted context omitted.

keyboard warrior strikes again lol. Most people would be thrilled to even be a small contributor in a tech initiative like this. call it what you want, your comment is just poor taste.

When Google did this with the recent Gemini paper, no one had any problem with calling it out as credential stuffing, but when Deepseek does it, it’s glorious unity and camaraderie.

Being the originator of this thread, I hold the same opinions about the Gemini paper from DeepMind, I see team spirit over cred stuffing

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#345

Even if you think this particular team cheated, the idea that nobody will find ways of making training more efficient seems silly - these huge datacenter investments for purely AI will IMHO seem very short sighted in 10 years

More like three years. Even in the best case the retained value curve of GPUs is absolutely terrible. Most of these huge investments in GPUs are going to be massive losses.

I actually wonder if this is true in the long term regardless of any AI uses. I mean, GPUs are general-purpose parallel compute, and there are so many things you can throw at them that can be of interest, whether economic or otherwise. For example, you can use them to model nuclear reactions...

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#346
post #175

DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...

Meta is in full panic last I heard. They have amassed a collection of pseudo experts there to collect their checks. Yet, Zuck wants to keep burning money on mediocrity. I’ve yet to see anything of value in terms products out of Meta.

>They have amassed a collection of pseudo experts there to collect their checks

LLaMA was huge, Byte Latent Transformer looks promising.. absolutely no idea were you got this idea from.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#347
post #175

Earlier quoted context omitted.

Meta is in full panic last I heard. They have amassed a collection of pseudo experts there to collect their checks. Yet, Zuck wants to keep burning money on mediocrity. I’ve yet to see anything of value in terms products out of Meta.

DeepSeek was built on the foundations of public research, a major part of which is the Llama family of models. Prior to Llama open weights LLMs were considerably less performant; without Llama we might not have gotten Mistral, Qwen, or DeepSeek. This isn't meant to diminish DeepSeek's contributions, however: they've been doing great work on mixture of experts models and really pushing the community forward on that fr…

I never said Llama is mediocre. I said the teams they put together is full of people chasing money. And the billions Meta is burning is going straight to mediocrity. They’re bloated. And we know exactly why Meta is doing this and it’s not because they have some grand scheme to build up AI. It’s to keep these people away from their competition. Same with billions in GPU spend. They want to suck up resources away from competition. That’s their entire plan. Do you really think Zuck has any clue about AI? He was never serious and instead built wonky VR prototypes.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#349

Earlier quoted context omitted.

The $500B is just an aspirational figure they hope to spend on data centers to run AI models, such as GPT-o1 and its successors, that have already been developed. If you want to compare the DeepSeek-R development costs to anything, you should be comparing it to what it cost OpenAI to develop GPT-o1 (not what they plan to spend to run it), but both numbers are somewhat irrelevant since they both build upon prior resea…

Thinking of the $500B as only an aspirational number is wrong. It’s true that the specific Stargate investment isn’t fully invested yet, but that’s hardly the only money being spent on AI development. The existing hyperscalers have already sunk ungodly amounts of money into literally hundreds of new data centers, millions of GPUs to fill them, chip manufacturing facilities, and even power plants with the impression t…

/Literally hundreds of billions of dollars spent already on hardware that’s already half (or fully) built, and isn’t easily repurposed./

It's just data centers full of devices optimized for fast linear algebra, right? These are extremely repurposeable.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#350
So is GRPO that much better because it ascribes feedback to a whole tight band of ‘quality’ ranges of on-policy answers while the band tends towards improvement in the aggregate, or is it just faster algorithm = more updates for a given training duration?
Post reply on HN