Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

591–600 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#591

Earlier quoted context omitted.

I must be missing something, but I tried Deepseek R1 via Kagi assistant and IMO it doesn't even come close to Claude? I don't get the hype at all? What am I doing wrong? And of course if you ask it anything related to the CCP it will suddenly turn into a Pinokkio simulator.

I haven't tried kagi assistant, but try it at deepseek.com. All models at this point have various politically motivated filters. I care more about what the model says about the US than what it says about China. Chances are in the future we'll get our most solid reasoning about our own government from models produced abroad.

> All models at this point have various politically motivated filters.

Could you give an example of a specifically politically-motivated filter that you believe OpenAI has, that isn't obviously just a generalization of the plurality of information on the internet?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#592

Neither of the deepseek models are on Groq yet, but when/if they are, that combination makes so much sense. A high quality open reasoning model, but you compensate for the slow inference of reasoning models with fast ASICs.

We are going to see it happen without something like next generation Groq chips. IIUC Groq can't run actually large LMs, the largest they offer is 70B LLaMA. DeepSeek-R1 is 671B.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#593
I've been comparing R1 to O1 and O1-pro, mostly in coding, refactoring and understanding of open source code.

I can say that R1 is on par with O1. But not as deep and capable as O1-pro. R1 is also a lot more useful than Sonnete. I actually haven't used Sonnete in awhile.

R1 is also comparable to the Gemini Flash Thinking 2.0 model, but in coding I feel like R1 gives me code that works without too much tweaking.

I often give entire open-source project's codebase (or big part of code) to all of them and ask the same question - like add a plugin, or fix xyz, etc. O1-pro is still a clear and expensive winner. But if I were to choose the second best, I would say R1.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#594

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

I don't get it. I like DeepSeek, because I can turn on Search button. Turning on Deepthink R1 makes the results as bad as Perplexity. The results make me feel like they used parallel construction, and that the straightforward replies would have actually had some value.

Claude Sonnet 3."6" may be limited in rare situations, but its personality really makes the responses outperform everything else when you're trying to take a deep dive into a subject where you previously knew nothing.

I think that the "thinking" part is a fiction, but it would be pretty cool if it gave you the thought process, and you could edit it. Often with these reasoning models like DeepSeek R1, the overview of the research strategy is nuts for the problem domain.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#595

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

I must be missing something, but I tried Deepseek R1 via Kagi assistant and IMO it doesn't even come close to Claude? I don't get the hype at all? What am I doing wrong? And of course if you ask it anything related to the CCP it will suddenly turn into a Pinokkio simulator.

Same here.

Following all the hype I tried it on my usual tasks (coding, image prompting...) and all I got was extra-verbose content with lower quality.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#596

Earlier quoted context omitted.

The word you're looking for is copyright enfrignment. That's the secret sause that every good model uses.

Humanity keeps running into copyright issues with every major leap in IT technology (photocopiers, tape cassettes, personal computers, internet, and now AI). I think it's about time for humanity to rethink their take on the unnatural restriction of information. I personally hope that countries recognize copyright and patents for what they really are and abolish them. Countries that refuse to do so can play catch up.

Since all kinds of companies are getting a lot of money from the generative AI business, I think they can handle being sued for plagiarism if thats the content they produce.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#597

I've been comparing R1 to O1 and O1-pro, mostly in coding, refactoring and understanding of open source code. I can say that R1 is on par with O1. But not as deep and capable as O1-pro. R1 is also a lot more useful than Sonnete. I actually haven't used Sonnete in awhile. R1 is also comparable to the Gemini Flash Thinking 2.0 model, but in coding I feel like R1 gives me code that works without too much tweaking. I oft…

How do you pass these models code bases?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#599
post #493

Earlier quoted context omitted.

From my casual read, right now everyone is on reputation tarnishing tirade, like spamming “Chinese stealing data! Definitely lying about everything! API can’t be this cheap!”. If that doesn’t go through well, I’m assuming lobbyism will start for import controls, which is very stupid. I have no idea how they can recover from it, if DeepSeek’s product is what they’re advertising.

Funny, everything I see (not actively looking for DeepSeek related content) is absolutely raving about it and talking about it destroying OpenAI (random YouTube thumbnails, most comments in this thread, even CNBC headlines). If DeepSeek's claims are accurate, then they themselves will be obsolete within a year, because the cost to develop models like this has dropped dramatically. There are going to be a lot of teams…

> If DeepSeek's claims are accurate, then they themselves will be obsolete within a year, because the cost to develop models like this has dropped dramatically. There are going to be a lot of teams with a lot of hardware resources with a lot of motivation to reproduce and iterate from here.

That would be an amazing outcome. For a while I was seriously worried about the possibility that if the trend of way more compute -> more AI breakthroughs continued, eventually AGI would be attained and exclusively controlled by a few people like Sam Altman who have trillions of $$$ to spend, and we’d all be replaced and live on whatever Sam-approved allowance.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#600

For context: R1 is a reasoning model based on V3. DeepSeek has claimed that GPU costs to train V3 (given prevailing rents) were about $5M. The true costs and implications of V3 are discussed here: https://www.interconnects.ai/p/deepseek-v3-and-the-actual-co...

This is great context for the cost claim. Which turns out only to be technically true when looking at the final run.
Post reply on HN