Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

501–510 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#501
post #472
post #459

Earlier quoted context omitted.

It will run faster than you can read on a MacBook Pro with 192GB.

You can only run a distilled model. They're quite good but not nearly as good as the full thing. As for as fast as you can read, depends on the distilled size. I have a mac mini 64 GB Ram. The 32 GB models are quite slow. 14B and lower are very very fast.

M4 or M4 Pro?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#502

Earlier quoted context omitted.

Curious why you have to qualify this with a “no fan of the CCP” prefix. From the outset, this is just a private organization and its links to CCP aren’t any different than, say, Foxconn’s or DJI’s or any of the countless Chinese manufacturers and businesses You don’t invoke “I’m no fan of the CCP” before opening TikTok or buying a DJI drone or a BYD car. Then why this, because I’ve seen the same line repeated everywh…

Any Chinese company above 500 employees requires a CCP representative on the board.

This is just an unfair clause set up to solve the employment problem of people within the system, to play a supervisory role and prevent companies from doing evil. In reality, it has little effect, and they still have to abide by the law.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#504

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

The nVidia market price could also be questionable considering how much cheaper DS is to run.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#505

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

There has never been much secret sauce in the model itself. The secret sauce or competitive advantage has always been in the engineering that goes into the data collection, model training infrastructure, and lifecycle/debugging management of model training. As well as in the access to GPUs.

Yeah, with Deepseek the barrier to entry has become significantly lower now. That's good, and hopefully more competition will come. But it's not like it's a fundamental change of where the secret sauce is.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#506
post #269

Earlier quoted context omitted.

Just commission the Chinese and make it 10X bigger then. In the case of the AI, they appear to commission Sam Altman and Larry Ellison.

The US has tried to commission Japan for that before. Japan gave up because we wouldn't do anything they asked and went to Morocco.

It was France:

https://www.businessinsider.com/french-california-high-speed...

Doubly delicious since the French have a long and not very nice colonial history in North Africa, sowing long-lasting suspicion and grudges, and still found it easier to operate there.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#507
post #504

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

The nVidia market price could also be questionable considering how much cheaper DS is to run.

It should be. I think AMD has left a lot on the table with respect to competing in the space (probably to the point of executive negligence) and the new US laws will help create several new Chinese competitors. NVIDIA probably has a bit of time left as the market leader, but it's really due mostly to luck.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#508
post #40

Question about the rule-based rewards (correctness and format) mentioned in the paper: Does the raw base model just expected “stumble upon“ a correct answer /correct format to get a reward and start the learning process? Are there any more details about the reward modelling?

The prompt in table 1 makes it very likely that the model will use the correct format. The pretrained model is pretty good so it only needs to stumble upon a correct answer every once in a while to start making progress. Some additional details in the Shao et al, 2024 paper.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#509

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

There has never been much secret sauce in the model itself. The secret sauce or competitive advantage has always been in the engineering that goes into the data collection, model training infrastructure, and lifecycle/debugging management of model training. As well as in the access to GPUs. Yeah, with Deepseek the barrier to entry has become significantly lower now. That's good, and hopefully more competition will co…

The word you're looking for is copyright enfrignment.

That's the secret sause that every good model uses.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#510
post #504

Earlier quoted context omitted.

The nVidia market price could also be questionable considering how much cheaper DS is to run.

It should be. I think AMD has left a lot on the table with respect to competing in the space (probably to the point of executive negligence) and the new US laws will help create several new Chinese competitors. NVIDIA probably has a bit of time left as the market leader, but it's really due mostly to luck.

As we have seen here it won't be a Western company that saves us from the dominant monopoly.

Xi Jinping, you're our only hope.

Post reply on HN