Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

61–70 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#61

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

Why do americans think china is like a hivemind controlled by an omnisicient Xi, making strategic moves to undermine them? Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x?

SAY WHAT?

Do you want an Internet without conspiracy theories?

Where have you been living for the last decades?

/s

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#62
The poor readability bit is quite interesting to me. While the model does develop some kind of reasoning abilities, we have no idea what the model is doing to convince itself about the answer. These could be signs of non-verbal reasoning, like visualizing things and such. Who knows if the model hasn't invented genuinely novel things when solving the hardest questions? And could the model even come up with qualitatively different and "non human" reasoning processes? What would that even look like?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#63

Earlier quoted context omitted.

Why do americans think china is like a hivemind controlled by an omnisicient Xi, making strategic moves to undermine them? Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x?

> Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x They literally published all their methodology. It's nothing groundbreaking, just western labs seem slow to adopt new research. Mixture of experts, key-value cache compression, multi-token prediction, 2/3 of these weren't invented by DeepSeek. They did invent a new hardware-aware distributed training approach for mixture…

"nothing groundbreaking"

It's extremely cheap, efficient and kicks the ass of the leader of the market, while being under sanctions with AI hardware.

Most of all, can be downloaded for free, can be uncensored, and usable offline.

China is really good at tech, it has beautiful landscapes, etc. It has its own political system, but to be fair, in some way it's all our future.

A bit of a dystopian future, like it was in 1984.

But the tech folks there are really really talented, it's long time that China switched from producing for the Western clients, to direct-sell to the Western clients.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#64

Earlier quoted context omitted.

Why do americans think china is like a hivemind controlled by an omnisicient Xi, making strategic moves to undermine them? Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x?

> Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x They literally published all their methodology. It's nothing groundbreaking, just western labs seem slow to adopt new research. Mixture of experts, key-value cache compression, multi-token prediction, 2/3 of these weren't invented by DeepSeek. They did invent a new hardware-aware distributed training approach for mixture…

That's what they claim at least in the paper but that particular claim is not verifiable. The HAI-LLM framework they reference in the paper is not open sourced and it seems they have no plans to.

Additionally there are claims, such as those by Scale AI CEO Alexandr Wang on CNBC 1/23/2025 time segment below, that DeepSeek has 50,000 H100s that "they can't talk about" due to economic sanctions (implying they likely got by avoiding them somehow when restrictions were looser). His assessment is that they will be more limited moving forward.

https://youtu.be/x9Ekl9Izd38?t=178

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#65

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

I don't think we were wrong to look at this as a commodity problem and ask how many widgets we need. Most people will still get their access to this technology through cloud services and nothing in this paper changes the calculations for inference compute demand. I still expect inference compute demand to be massive and distilled models aren't going to cut it for most agentic use cases.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#66

Earlier quoted context omitted.

Why do americans think china is like a hivemind controlled by an omnisicient Xi, making strategic moves to undermine them? Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x?

China is actually just one person (Xi) acting in perfect unison and its purpose is not to benefit its own people, but solely to undermine the West.

This explains so much. It’s just malice, then? Or some demonic force of evil? What does Occam’s razor suggest?

Oh dear

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#67
I've always been leery about outrageous GPU investments, at some point I'll dig through and find my prior comments where I've said as much to that effect.

The CEOs, upper management, and governments derive their importance on how much money they can spend - AI gave them the opportunity for them to confidently say that if you give me $X I can deliver Y and they turn around and give that money to NVidia. The problem was reduced to a simple function of raising money and spending that money making them the most importance central figure. ML researchers are very much secondary to securing funding. Since these people compete with each other in importance they strived for larger dollar figures - a modern dick waving competition. Those of us who lobbied for efficiency were sidelined as we were a threat. It was seen as potentially making the CEO look bad and encroaching in on their importance. If the task can be done for cheap by smart people then that severely undermines the CEOs value proposition.

With the general financialization of the economy the wealth effect of the increase in the cost of goods increases wealth by a greater amount than the increase in cost of goods - so that if the cost of housing goes up more people can afford them. This financialization is a one way ratchet. It appears that the US economy was looking forward to blowing another bubble and now that bubble has been popped in its infancy. I think the slowness of the popping of this bubble underscores how little the major players know about what has just happened - I could be wrong about that but I don't know how yet.

Edit: "[big companies] would much rather spend huge amounts of money on chips than hire a competent researcher who might tell them that they didn’t really need to waste so much money." (https://news.ycombinator.com/item?id=39483092 11 months ago)

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#68

Earlier quoted context omitted.

> Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x They literally published all their methodology. It's nothing groundbreaking, just western labs seem slow to adopt new research. Mixture of experts, key-value cache compression, multi-token prediction, 2/3 of these weren't invented by DeepSeek. They did invent a new hardware-aware distributed training approach for mixture…

That's what they claim at least in the paper but that particular claim is not verifiable. The HAI-LLM framework they reference in the paper is not open sourced and it seems they have no plans to. Additionally there are claims, such as those by Scale AI CEO Alexandr Wang on CNBC 1/23/2025 time segment below, that DeepSeek has 50,000 H100s that "they can't talk about" due to economic sanctions (implying they likely got…

It's amazing how different the standards are here. Deepseek's released their weights under a real open source license and published a paper with their work which now has independent reproductions.

OpenAI literally haven't said a thing about how O1 even works.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#69

Genuinly curious, what is everyone using reasoning models for? (R1/o1/o3)

We've been seeing success using it for LLM-as-a-judge tasks.

We set up an evaluation criteria and used o1 to evaluate the quality of the prod model, where the outputs are subjective, like creative writing or explaining code.

It's also useful for developing really good few-shot examples. We'll get o1 to generate multiple examples in different styles, then we'll have humans go through and pick the ones they like best, which we use as few-shot examples for the cheaper, faster prod model.

Finally, for some study I'm doing, I'll use it to grade my assignments before I hand them in. If I get a 7/10 from o1, I'll ask it to suggest the minimal changes I could make to take it to 10/10. Then, I'll make the changes and get it to regrade the paper.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#70

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

> The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value.

I do not quite follow. GPU compute is mostly spent in inference, as training is a one time cost. And these chain of thought style models work by scaling up inference time compute, no?

So proliferation of these types of models would portend in increase in demand for GPUs?

Post reply on HN