Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

671–680 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#671

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

Most people I talked with don't grasp how big of an event this is. I consider is almost as similar to as what early version of linux did to OS ecosystem.

That is probably because they did not try the model yet. I tried and was stunned. It's not better yet in all areas, but where is better, is so much better than Claude or anything from OpenAI.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#672
post #182

DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...

Correct me if I'm wrong but if Chinese can produce the same quality at %99 discount, then the supposed $500B investment is actually worth $5B. Isn't that the kind wrong investment that can break nations? Edit: Just to clarify, I don't imply that this is public money to be spent. It will commission $500B worth of human and material resources for 5 years that can be much more productive if used for something else - i.e…

When we move to continuously running agents, rather than query-response models, we're going to need a lot more compute.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#673
post #671

Earlier quoted context omitted.

Most people I talked with don't grasp how big of an event this is. I consider is almost as similar to as what early version of linux did to OS ecosystem.

That is probably because they did not try the model yet. I tried and was stunned. It's not better yet in all areas, but where is better, is so much better than Claude or anything from OpenAI.

Plus, the speed at which it replies is amazing too. Claude/Chatgpt now seem like inefficient inference engines compared to it.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#674
post #605

Earlier quoted context omitted.

False equivalency. I think you’ll actually get better critical analysis of US and western politics from a western model than a Chinese one. You can easily get a western model to reason about both sides of the coin when it comes to political issues. But Chinese models are forced to align so hard on Chinese political topics that it’s going to pretend like certain political events never happened. E.g try getting them to…

GPT4 is also full of ideology, but of course the type you probably grew up with, so harder to see. (No offense intended, this is just the way ideology works). Try for example to persuade GPT to argue that the workers doing data labeling in Kenya should be better compensated relative to the programmers in SF, as the work they do is both critical for good data for training and often very gruesome, with many workers get…

If you've forced OpenAI to pay Kenyans as much as Americans, then OpenAI simply would stop hiring Kenyans. Beware of the unintended consequences of your ideological narrative.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#675
post #535

Earlier quoted context omitted.

Nah, this just means training isn’t the advantage. There’s plenty to be had by focusing on inference. It’s like saying apple is dead because back in 1987 there was a cheaper and faster PC offshore. I sure hope so otherwise this is a pretty big moment to question life goals.

> saying apple is dead because back in 1987 there was a cheaper and faster PC offshore What Apple did was build a luxury brand and I don't see that happening with LLMs. When it comes to luxury, you really can't compete with price.

Apple isn’t a luxury brand in the normal sense, it’s odd that people think this because they’re more expensive. They’re not the technical equivalent of Prada or Rolex etc. Apple’s ecosystem cohesion and still unmatched UX (still flawed) is a real value-add that normal luxury brands don’t have.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#676
post #594

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

I don't get it. I like DeepSeek, because I can turn on Search button. Turning on Deepthink R1 makes the results as bad as Perplexity. The results make me feel like they used parallel construction, and that the straightforward replies would have actually had some value. Claude Sonnet 3."6" may be limited in rare situations, but its personality really makes the responses outperform everything else when you're trying to…

IMO the deep think button works wonders.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#677

Earlier quoted context omitted.

I asked Chatgpt: how many civilians Israel killed in Gaza. Please provide a rough estimate. As of January 2025, the conflict between Israel and Hamas has resulted in significant civilian casualties in the Gaza Strip. According to reports from the United Nations Office for the Coordination of Humanitarian Affairs (OCHA), approximately 7,000 Palestinian civilians have been killed since the escalation began in October 2…

Isn't the real number around 46,000 people, though?

At least according to the OCHA you're right . Though there's also a dashboard which shows around 7k for the entire Israel Palestine conflict since 2008. Maybe it got confused by the conflicting info on OCHA's website.

https://www.ochaopt.org/data/casualties

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#678
post #573

Tangentially the model seems to be trained in an unprofessional mode, using many filler words like 'okay' 'hmm' maybe it's done to sound cute or approachable but I find it highly annoying or is this how the model learns to talk through reinforcement learning and they didn't fix it with supervised reinforcement learning

I’m sure I’ve seen this technique in chain of thought before, where the model is instructed about certain patterns of thinking: “Hmm, that doesn’t seem quite right”, “Okay, now what?”, “But…”, to help it identify when reasoning is going down the wrong path. Which apparently increased the accuracy. It’s possible these filler words aren’t unprofessional but are in fact useful. If anyone can find a source for that I’d l…

I remember reading a paper that showed that giving models even a a few filler tokens before requiring a single phrase/word/number answer significantly increasee accuracy. This is probably similar.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#679
post #311
post #146

Larry Ellison is 80. Masayoshi Son is 67. Both have said that anti-aging and eternal life is one of their main goals with investing toward ASI. For them it's worth it to use their own wealth and rally the industry to invest $500 billion in GPUs if that means they will get to ASI 5 years faster and ask the ASI to give them eternal life.

Side note: I’ve read enough sci-fi to know that letting rich people live much longer than not rich is a recipe for a dystopian disaster. The world needs incompetent heirs to waste most of their inheritance, otherwise the civilization collapses to some kind of feudal nightmare.

Yeah imagine progress without the planck quote "science progresses one funeral at a time"

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#680

Earlier quoted context omitted.

I haven't tried kagi assistant, but try it at deepseek.com. All models at this point have various politically motivated filters. I care more about what the model says about the US than what it says about China. Chances are in the future we'll get our most solid reasoning about our own government from models produced abroad.

> All models at this point have various politically motivated filters. Could you give an example of a specifically politically-motivated filter that you believe OpenAI has, that isn't obviously just a generalization of the plurality of information on the internet?

Gemini models won't touch a lot of things that are remotely political in nature. One time I tried to use GPT-4o to verify some claims I read on the internet and it was very outspoken about issues relating to alleged election fraud, to the point where it really got in the way.

I generally find it unhelpful whaen models produce boilerplate meant to couch the response in any way.

Post reply on HN