Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

601–610 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#601

Earlier quoted context omitted.

The word you're looking for is copyright enfrignment. That's the secret sause that every good model uses.

It will be interesting if a significant jurisdiction's copyright law is some day changed to treat LLM training as copying. In a lot of places, previous behaviour can't be retroactively outlawed[1]. So older LLMs will be much more capable than post-change ones. [1] https://en.wikipedia.org/wiki/Ex_post_facto_law

Even if you can't be punished retroactively for previous behavior, continuing to benefit from it can be outlawed. In other words, it would be compatible from a legal perspective to ban the use of LLMs that were trained in violation of copyright law.

Given the political landscape I doubt that's going to happen, though.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#602
post #9

we've been tracking the deepseek threads extensively in LS. related reads: - i consider the deepseek v3 paper required preread https://github.com/deepseek-ai/DeepSeek-V3 - R1 + Sonnet > R1 or O1 or R1+R1 or O1+Sonnet or any other combo https://aider.chat/2025/01/24/r1-sonnet.html - independent repros: 1) https://hkust-nlp.notion.site/simplerl-reason 2) https://buttondown.com/ainews/archive/ainews-tinyzero-reprod... 3…

In the context of tracking DeepSeek threads, "LS" could plausibly stand for: 1. *Log System/Server*: A platform for storing or analyzing logs related to DeepSeek's operations or interactions. 2. *Lab/Research Server*: An internal environment for testing, monitoring, or managing AI/thread data. 3. *Liaison Service*: A team or interface coordinating between departments or external partners. 4. *Local Storage*: A reposi…

Latent space

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#603

I've been comparing R1 to O1 and O1-pro, mostly in coding, refactoring and understanding of open source code. I can say that R1 is on par with O1. But not as deep and capable as O1-pro. R1 is also a lot more useful than Sonnete. I actually haven't used Sonnete in awhile. R1 is also comparable to the Gemini Flash Thinking 2.0 model, but in coding I feel like R1 gives me code that works without too much tweaking. I oft…

At this point, it's a function of how many thinking tokens can a model generate. (when it comes to o1 and r1). o3 is likely going to be superior because they used the training data generated from o1 (amongst other things). o1-pro has a longer "thinking" token length, so it comes out as better. Same goes with o1 and API where you can control the thinking length. I have not seen the implementation for r1 api as such, but if they provide that option, the output could be even better.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#604

Earlier quoted context omitted.

Those are not just-throw-money problems. Usually these tropes are limited to instagram comments. Surprised to see it here.

I know, it was simply to show the absurdity of committing $500B to marginally improving next token predictors.

True. I think there is some posturing involved in the 500b number as well.

Either that or its an excuse for everyone involved to inflate the prices.

Hopefully the datacenters are useful for other stuff as well. But also I saw a FT report that it's going to be exclusive to openai?

Also as I understand it these types of deals are usually all done with speculative assets. And many think the current AI investments are a bubble waiting to pop.

So it will still remain true that if jack falls down and breaks his crown, jill will be tumbling after.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#605

Earlier quoted context omitted.

I must be missing something, but I tried Deepseek R1 via Kagi assistant and IMO it doesn't even come close to Claude? I don't get the hype at all? What am I doing wrong? And of course if you ask it anything related to the CCP it will suddenly turn into a Pinokkio simulator.

I haven't tried kagi assistant, but try it at deepseek.com. All models at this point have various politically motivated filters. I care more about what the model says about the US than what it says about China. Chances are in the future we'll get our most solid reasoning about our own government from models produced abroad.

False equivalency. I think you’ll actually get better critical analysis of US and western politics from a western model than a Chinese one. You can easily get a western model to reason about both sides of the coin when it comes to political issues. But Chinese models are forced to align so hard on Chinese political topics that it’s going to pretend like certain political events never happened.

E.g try getting them to talk in a critical way about “the trail of tears” and “tiananmen square”

It could be interesting to challenge these models on something like the rights of Hawaiian people and the possibility of Hawaii independence. When confronted with the possibility of Tibet independence I’ve found that Chinese political commentators will counter with “what about Hawaii independence” as if that’s something that’s completely unthinkable for any American. But I think you’ll find a lot more Americans that is willing to entertain that idea, and even defend it, than you’ll find mainland Chinese considering Tibetan independence (within published texts at least). So I’m sceptical about a Chinese models ability to accurately tackle the question of the rights of a minority population within an empire, in a fully consistent way.

Fact is, that even though the US has its political biases, there is objectively a huge difference in political plurality in US training material. Hell, it may even have “Xi Jinping thought” in there

And I think it’s fair to say that a model that has more plurality in its political training data will be much more capable and useful in analysing political matters.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#606
post #592

Neither of the deepseek models are on Groq yet, but when/if they are, that combination makes so much sense. A high quality open reasoning model, but you compensate for the slow inference of reasoning models with fast ASICs.

We are going to see it happen without something like next generation Groq chips. IIUC Groq can't run actually large LMs, the largest they offer is 70B LLaMA. DeepSeek-R1 is 671B.

Aha, for some reason I thought they provided full-size Llama through some bundling of multiple chips. Fair enough then, anyway long term I feel like providers running powerful open models on purpose built inference ASICs will be really awesome.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#607
post #119

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

Seeing what china is doing to the car market, I give it 5 years for China to do to the AI/GPU market to do the same. This will be good. Nvidia/OpenAI monopoly is bad for everyone. More competition will be welcome.

AI sure, which is good, as I'd rather not have giant companies in the US monopolizing it. If they open source it and undercut OpenAI etc all the better

GPU: nope, that would take much longer, Nvidia/ASML/TSMC is too far ahead

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#608
post #504

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

The nVidia market price could also be questionable considering how much cheaper DS is to run.

Deepseek has thousands of Nvidia GPUs, though.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#609

Earlier quoted context omitted.

I must be missing something, but I tried Deepseek R1 via Kagi assistant and IMO it doesn't even come close to Claude? I don't get the hype at all? What am I doing wrong? And of course if you ask it anything related to the CCP it will suddenly turn into a Pinokkio simulator.

I told it to write its autobiography via DeepSeek chat and it told me it _was_ Claude. Which is a little suspicious.

One report is an anecdote, but I wouldn't be surprised if we heard more of this. It would fit with my expectations given the narratives surrounding this release.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#610

DeepSeek V3 came in the perfect time, precisely when Claude Sonnet turned into crap and barely allows me to complete something without me hitting some unexpected constraints. Idk, what their plans is and if their strategy is to undercut the competitors but for me, this is a huge benefit. I received 10$ free credits and have been using Deepseeks api a lot, yet, I have barely burned a single dollar, their pricing are t…

Can you tell me more about how Claude Sonnet went bad for you? I've been using the free version pretty happily, and felt I was about to upgrade to paid any day now (well, at least before the new DeepSeek).

I should’ve maybe been more explicit, it’s Claudes service that I think sucks atm, not their model.

It feels like the free quota has been lowered much more than previously, and I have been using it since it was available to EU.

I can’t count how many times I’ve started a conversation and after a couple of messages I get ”unexpected constrain (yada yada)”. It is either that or I get a notification saying ”defaulting to Haiku because of high demand”.

I don’t even have long conversations because I am aware of how longer conversations can use up the free quota faster, my strategy is to start a new conversation with a little context as soon as I’ve completed the task.

I’ve had thoughts about paying for a subscription because how much I enjoy Sonnet 3.5, but it is too expensive for me and I don’t use it that much to pay 20$ monthly.

My suspicion is that Claude has gotten very popular since the beginning of last year and now Anthropic have hit their maximum capacity.

This is why I said DeepSeek came in like a savior, it performs close to Claude but for pennies, it’s amazing!

Post reply on HN