Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

71–80 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#71

Earlier quoted context omitted.

Why do americans think china is like a hivemind controlled by an omnisicient Xi, making strategic moves to undermine them? Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x?

> Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x They literally published all their methodology. It's nothing groundbreaking, just western labs seem slow to adopt new research. Mixture of experts, key-value cache compression, multi-token prediction, 2/3 of these weren't invented by DeepSeek. They did invent a new hardware-aware distributed training approach for mixture…

The U.S. firms let everyone skeptical go the second they had a marketable proof of concept, and replaced them with smart, optimistic, uncritical marketing people who no longer know how to push the cutting edge.

Maybe we don't need momentum right now and we can cut the engines.

Oh, you know how to develop novel systems for training and inference? Well, maybe you can find 4 people who also can do that by breathing through the H.R. drinking straw, and that's what you do now.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#72
post #9

we've been tracking the deepseek threads extensively in LS. related reads: - i consider the deepseek v3 paper required preread https://github.com/deepseek-ai/DeepSeek-V3 - R1 + Sonnet > R1 or O1 or R1+R1 or O1+Sonnet or any other combo https://aider.chat/2025/01/24/r1-sonnet.html - independent repros: 1) https://hkust-nlp.notion.site/simplerl-reason 2) https://buttondown.com/ainews/archive/ainews-tinyzero-reprod... 3…

Hugging Face is reproducing R1 in public.

https://x.com/_lewtun/status/1883142636820676965

https://github.com/huggingface/open-r1

Hugging Face Journal Club - DeepSeek R1 https://www.youtube.com/watch?v=1xDVbu-WaFo

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#73

Earlier quoted context omitted.

China is actually just one person (Xi) acting in perfect unison and its purpose is not to benefit its own people, but solely to undermine the West.

This explains so much. It’s just malice, then? Or some demonic force of evil? What does Occam’s razor suggest? Oh dear

Always attribute to malice what can’t be explained by mere stupidity. ;)

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#75

Earlier quoted context omitted.

Why do americans think china is like a hivemind controlled by an omnisicient Xi, making strategic moves to undermine them? Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x?

China is actually just one person (Xi) acting in perfect unison and its purpose is not to benefit its own people, but solely to undermine the West.

If China is undermining the West by lifting up humanity, for free, while ProprietaryAI continues to use closed source AI for censorship and control, then go team China.

There's something wrong with the West's ethos if we think contributing significantly to the progress of humanity is malicious. The West's sickness is our own fault; we should take responsibility for our own disease, look critically to understand its root, and take appropriate cures, even if radical, to resolve our ailments.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#76
post #63

Earlier quoted context omitted.

> Is it really that unlikely that a lab of genius engineers found a way to improve efficiency 10x They literally published all their methodology. It's nothing groundbreaking, just western labs seem slow to adopt new research. Mixture of experts, key-value cache compression, multi-token prediction, 2/3 of these weren't invented by DeepSeek. They did invent a new hardware-aware distributed training approach for mixture…

"nothing groundbreaking" It's extremely cheap, efficient and kicks the ass of the leader of the market, while being under sanctions with AI hardware. Most of all, can be downloaded for free, can be uncensored, and usable offline. China is really good at tech, it has beautiful landscapes, etc. It has its own political system, but to be fair, in some way it's all our future. A bit of a dystopian future, like it was in…

[dead]

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#77
post #40

Question about the rule-based rewards (correctness and format) mentioned in the paper: Does the raw base model just expected “stumble upon“ a correct answer /correct format to get a reward and start the learning process? Are there any more details about the reward modelling?

Good question.

When BF Skinner used to train his pigeons, he’d initially reinforce any tiny movement that at least went in the right direction. For the exact reasons you mentioned.

For example, instead of waiting for the pigeon to peck the lever directly (which it might not do for many hours), he’d give reinforcement if the pigeon so much as turned its head towards the lever. Over time, he’d raise the bar. Until, eventually, only clear lever pecks would receive reinforcement.

I don’t know if they’re doing something like that here. But it would be smart.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#78
post #26

Earlier quoted context omitted.

CEO of Scale said Deepseek is lying and actually has a 50k GPU cluster. He said they lied in the paper because technically they aren't supposed to have them due to export laws. I feel like this is very likely. They obvious did some great breakthroughs, but I doubt they were able to train on so much less hardware.

I would think the CEO of an American AI company has every reason to neg and downplay foreign competition... And since it's a businessperson they're going to make it sound as cute and innocuous as possible

It's hard to tell if they're telling the truth about the number of GPUs they have. They open sourced the model and the inference is much more efficient than the best American models so it's not implausible that the training was also much more efficient.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#79
How can openai justify their $200/mo subscriptions if a model like this exists at an incredibly low price point? Operator?

I've been impressed in my brief personal testing and the model ranks very highly across most benchmarks (when controlled for style it's tied number one on lmarena).

It's also hilarious that openai explicitly prevented users from seeing the CoT tokens on the o1 model (which you still pay for btw) to avoid a situation where someone trained on that output. Turns out it made no difference lmao.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#80
There seems to be a print out of "reasoning". Is that some new breaktheough thing? Really impressive.

E.g. I tried to make it guess my daughter's name and I could only answer yes or no and the first 5 questions where very convincing but then it lost track and started to randomly guess names one by one.

edit: Nagging it to narrow it down and give a language group hint made it solve it. Ye, well, it can do Akinator.

Post reply on HN