Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

531–540 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#531

Over 100 authors on arxiv and published under the team name, that's how you recognize everyone and build comradery. I bet morale is high over there

Except now you end up with folks who probably ran some analysis or submitted some code changes getting thousands of citations on Google Scholar for DeepSeek.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#532

Earlier quoted context omitted.

The chain of thought is super useful in so many ways, helping me: (1) learn, way beyond the final answer itself, (2) refine my prompt, whether factually or stylistically, (3) understand or determine my confidence in the answer.

do you have any resources related to these???

What do you mean? I was referring to just the chain of thought you see when the "DeepThink (R1)" button is enabled. As someone who LOVES learning (as many of you too), R1 chain of thought is an infinite candy store.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#533

Earlier quoted context omitted.

I must be missing something, but I tried Deepseek R1 via Kagi assistant and IMO it doesn't even come close to Claude? I don't get the hype at all? What am I doing wrong? And of course if you ask it anything related to the CCP it will suddenly turn into a Pinokkio simulator.

I haven't tried kagi assistant, but try it at deepseek.com. All models at this point have various politically motivated filters. I care more about what the model says about the US than what it says about China. Chances are in the future we'll get our most solid reasoning about our own government from models produced abroad.

[deleted]

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#534

Earlier quoted context omitted.

There has never been much secret sauce in the model itself. The secret sauce or competitive advantage has always been in the engineering that goes into the data collection, model training infrastructure, and lifecycle/debugging management of model training. As well as in the access to GPUs. Yeah, with Deepseek the barrier to entry has become significantly lower now. That's good, and hopefully more competition will co…

I don't disagree, but the important point is that Deepseek showed that it's not just about CapEx, which is what the US firms were/are lining up to battle with. In my opinion there is something qualitatively better about Deepseek in spite of its small size, even compared to o1-pro, that suggests a door has been opened. GPUs are needed to rapidly iterate on ideas, train, evaluate, etc., but Deepseek has shown us that w…

How do you know the CCP didn’t just help out with lots of compute and then tell the companies to lie about how much it cost to train the model?

Reagan did the same with Star Wars, in order to throw the USSR into exactly the same kind of competition hysteria and try to bankrupt it. And USA today is very much in debt as it is… seems like a similar move:

https://www.nytimes.com/1993/08/18/us/lies-and-rigged-star-w...

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#535

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

Nah, this just means training isn’t the advantage. There’s plenty to be had by focusing on inference. It’s like saying apple is dead because back in 1987 there was a cheaper and faster PC offshore. I sure hope so otherwise this is a pretty big moment to question life goals.

> saying apple is dead because back in 1987 there was a cheaper and faster PC offshore

What Apple did was build a luxury brand and I don't see that happening with LLMs. When it comes to luxury, you really can't compete with price.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#536
Tangentially the model seems to be trained in an unprofessional mode, using many filler words like 'okay' 'hmm' maybe it's done to sound cute or approachable but I find it highly annoying

or is this how the model learns to talk through reinforcement learning and they didn't fix it with supervised reinforcement learning

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#539

Earlier quoted context omitted.

There has never been much secret sauce in the model itself. The secret sauce or competitive advantage has always been in the engineering that goes into the data collection, model training infrastructure, and lifecycle/debugging management of model training. As well as in the access to GPUs. Yeah, with Deepseek the barrier to entry has become significantly lower now. That's good, and hopefully more competition will co…

The word you're looking for is copyright enfrignment. That's the secret sause that every good model uses.

since all models are treating human knowledge as copyright free (as they should) no this is not at all what this new Chinese model is about

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#540

Earlier quoted context omitted.

The word you're looking for is copyright enfrignment. That's the secret sause that every good model uses.

since all models are treating human knowledge as copyright free (as they should) no this is not at all what this new Chinese model is about

Oh. Does that ethics framework also extend to art such as music, movies and software?

fires up BitTorrent

Post reply on HN