Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

521–530 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#521

Earlier quoted context omitted.

I don't disagree, but the important point is that Deepseek showed that it's not just about CapEx, which is what the US firms were/are lining up to battle with. In my opinion there is something qualitatively better about Deepseek in spite of its small size, even compared to o1-pro, that suggests a door has been opened. GPUs are needed to rapidly iterate on ideas, train, evaluate, etc., but Deepseek has shown us that w…

Let me qualify your statement... CapEx is what EXISTING US firms were/are lining up to battle with. With R1 as inspiration/imperative, many new US startups will emerge who will be very strong. Can you feel a bunch of talent in limbo startups pivoting/re-energized now?

> Can you feel a bunch of talent in limbo startups pivoting/re-energized now?

True! It certainly should be, as there is a lot less reason to hitch one's wagon to one of the few big firms that can afford nation state scale GPU compute.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#522

Earlier quoted context omitted.

As we have seen here it won't be a Western company that saves us from the dominant monopoly. Xi Jinping, you're our only hope.

If China really released a GPU competitive with the current generation of nvidia you can bet it'd be banned in the US like BYD and DJI.

Sad but likely true.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#523

Earlier quoted context omitted.

As we have seen here it won't be a Western company that saves us from the dominant monopoly. Xi Jinping, you're our only hope.

If China really released a GPU competitive with the current generation of nvidia you can bet it'd be banned in the US like BYD and DJI.

Ok but that leaves the rest of the world to China.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#524
post #478

Earlier quoted context omitted.

Can you tell me more about how Claude Sonnet went bad for you? I've been using the free version pretty happily, and felt I was about to upgrade to paid any day now (well, at least before the new DeepSeek).

It's not their model being bad, it's claude.ai having pretty low quota for even paid users. It looks like Anthropic doesn't have enough GPUs. It's not only claude.ai, they recently pushed back increasing API demand from Cursor too.

Interesting insight/possibility. I did see some capacity glitches with my Cursor recently. Overall, I like Anthropic (and ChatGPT); hopefully they continue to succeed.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#525
post #170

Earlier quoted context omitted.

Same thing happened to Google Gemini paper (1000+ authors) and it was described as big co promo culture (everyone wants credits). Interesting how narratives shift https://arxiv.org/abs/2403.05530

For me that sort of thing actually dilutes the prestige. If I'm interviewing someone, and they have "I was an author on this amazing paper!" on their resume, then if I open the paper and find 1k+ authors on it, at that point it's complete noise to me. I have absolutely no signal on their relative contributions vs. those of anyone else in the author list. At that point it's not really a publication, for all intents an…

That's how it works in most scientific fields. If you want more granularity, you check the order of the authors. Sometimes, they explaine in the paper who did what.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#526

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

Nah, this just means training isn’t the advantage. There’s plenty to be had by focusing on inference. It’s like saying apple is dead because back in 1987 there was a cheaper and faster PC offshore. I sure hope so otherwise this is a pretty big moment to question life goals.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#527

DeepSeek V3 came in the perfect time, precisely when Claude Sonnet turned into crap and barely allows me to complete something without me hitting some unexpected constraints. Idk, what their plans is and if their strategy is to undercut the competitors but for me, this is a huge benefit. I received 10$ free credits and have been using Deepseeks api a lot, yet, I have barely burned a single dollar, their pricing are t…

Can you tell me more about how Claude Sonnet went bad for you? I've been using the free version pretty happily, and felt I was about to upgrade to paid any day now (well, at least before the new DeepSeek).

I've been a paid Claude user almost since they offered it. IMO it works perfectly well still - I think people are getting into trouble running extremely long conversations and blowing their usage limit (which is not very clearly explained). With Claude Desktop it's always good practice to summarize and restart the conversation often.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#528

Earlier quoted context omitted.

If you ask it about the Tienanmen Square Massacre its "thought process" is very interesting.

[I typed something dumb while half asleep]

I'm not sure censorship or lack of it matters for most use cases. Why would businesses using LLM to speed up their processes, or a programmer using it to write code care about how accurately it answers to political questions?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#529
post #430

Earlier quoted context omitted.

I literally cannot see how OpenAI and Anthropic can justify their valuation given DeepSeek. In business, if you can provide twice the value at half the price, you will destroy the incumbent. Right now, DeepSeek is destroying on price and provides somewhat equivalent value compared to Sonnet. I still believe Sonnet is better, but I don't think it is 10 times better. Something else that DeepSeek can do, which I am not…

> I still believe Sonnet is better, but I don't think it is 10 times better. Sonnet doesn't need to be 10 times better. It just needs to be better enough such that the downstream task improves more than the additional cost. This is a much more reasonable hurdle. If you're able to improve the downstream performance of something that costs $500k/year by 1% then the additional cost of Sonnet just has to be less than $5k…

> But I don't think R1 is terminal for them.

I hope not, as I we need more competition.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#530

Earlier quoted context omitted.

Is it propaganda if it's true?

I haven't been to China since 2019, but it is pretty obvious that median quality of life is higher in the US. In China, as soon as you get out of Beijing-Shanghai-Guangdong cities you start seeing deep poverty, people in tiny apartments that are falling apart, eating meals in restaurants that are falling apart, and the truly poor are emaciated. Rural quality of life is much higher in the US.

Well, in the US you have millions of foreigners and blacks who live in utter poverty, and sustain the economy, just like the farmers in China.
Post reply on HN