Over 100 authors on arxiv and published under the team name, that's how you recognize everyone and build comradery. I bet morale is high over there
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
531–540 of 1001 posts
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#532Earlier quoted context omitted.
The chain of thought is super useful in so many ways, helping me: (1) learn, way beyond the final answer itself, (2) refine my prompt, whether factually or stylistically, (3) understand or determine my confidence in the answer.
do you have any resources related to these???
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#533Earlier quoted context omitted.
I must be missing something, but I tried Deepseek R1 via Kagi assistant and IMO it doesn't even come close to Claude? I don't get the hype at all? What am I doing wrong? And of course if you ask it anything related to the CCP it will suddenly turn into a Pinokkio simulator.
I haven't tried kagi assistant, but try it at deepseek.com. All models at this point have various politically motivated filters. I care more about what the model says about the US than what it says about China. Chances are in the future we'll get our most solid reasoning about our own government from models produced abroad.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#534Earlier quoted context omitted.
There has never been much secret sauce in the model itself. The secret sauce or competitive advantage has always been in the engineering that goes into the data collection, model training infrastructure, and lifecycle/debugging management of model training. As well as in the access to GPUs. Yeah, with Deepseek the barrier to entry has become significantly lower now. That's good, and hopefully more competition will co…
I don't disagree, but the important point is that Deepseek showed that it's not just about CapEx, which is what the US firms were/are lining up to battle with. In my opinion there is something qualitatively better about Deepseek in spite of its small size, even compared to o1-pro, that suggests a door has been opened. GPUs are needed to rapidly iterate on ideas, train, evaluate, etc., but Deepseek has shown us that w…
Reagan did the same with Star Wars, in order to throw the USSR into exactly the same kind of competition hysteria and try to bankrupt it. And USA today is very much in debt as it is… seems like a similar move:
https://www.nytimes.com/1993/08/18/us/lies-and-rigged-star-w...
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#535For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…
Nah, this just means training isn’t the advantage. There’s plenty to be had by focusing on inference. It’s like saying apple is dead because back in 1987 there was a cheaper and faster PC offshore. I sure hope so otherwise this is a pretty big moment to question life goals.
What Apple did was build a luxury brand and I don't see that happening with LLMs. When it comes to luxury, you really can't compete with price.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#536or is this how the model learns to talk through reinforcement learning and they didn't fix it with supervised reinforcement learning
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#537Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#538Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#539Earlier quoted context omitted.
There has never been much secret sauce in the model itself. The secret sauce or competitive advantage has always been in the engineering that goes into the data collection, model training infrastructure, and lifecycle/debugging management of model training. As well as in the access to GPUs. Yeah, with Deepseek the barrier to entry has become significantly lower now. That's good, and hopefully more competition will co…
The word you're looking for is copyright enfrignment. That's the secret sause that every good model uses.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#540Earlier quoted context omitted.
The word you're looking for is copyright enfrignment. That's the secret sause that every good model uses.
since all models are treating human knowledge as copyright free (as they should) no this is not at all what this new Chinese model is about
fires up BitTorrent