Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

651–660 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#651
post #605

Earlier quoted context omitted.

I haven't tried kagi assistant, but try it at deepseek.com. All models at this point have various politically motivated filters. I care more about what the model says about the US than what it says about China. Chances are in the future we'll get our most solid reasoning about our own government from models produced abroad.

False equivalency. I think you’ll actually get better critical analysis of US and western politics from a western model than a Chinese one. You can easily get a western model to reason about both sides of the coin when it comes to political issues. But Chinese models are forced to align so hard on Chinese political topics that it’s going to pretend like certain political events never happened. E.g try getting them to…

GPT4 is also full of ideology, but of course the type you probably grew up with, so harder to see. (No offense intended, this is just the way ideology works).

Try for example to persuade GPT to argue that the workers doing data labeling in Kenya should be better compensated relative to the programmers in SF, as the work they do is both critical for good data for training and often very gruesome, with many workers getting PTSD from all the horrible content they filter out.

I couldn't, about a year ago. The model always tried to argue in favor of the status quo because of market forces - which is, of course, axiomatic ideology.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#652
post #347

Earlier quoted context omitted.

DeepSeek was built on the foundations of public research, a major part of which is the Llama family of models. Prior to Llama open weights LLMs were considerably less performant; without Llama we might not have gotten Mistral, Qwen, or DeepSeek. This isn't meant to diminish DeepSeek's contributions, however: they've been doing great work on mixture of experts models and really pushing the community forward on that fr…

I never said Llama is mediocre. I said the teams they put together is full of people chasing money. And the billions Meta is burning is going straight to mediocrity. They’re bloated. And we know exactly why Meta is doing this and it’s not because they have some grand scheme to build up AI. It’s to keep these people away from their competition. Same with billions in GPU spend. They want to suck up resources away from…

In contrast to the Social Media industry (or word processors or mobile phones), the market for AI solutions seems not to have of an inherent moat or network effects which keep the users stuck in the market leader.

Rather with AI, capitalism seems working at its best with competitors to OpenAI building solutions which take market share and improve products. Zuck can try monopoly plays all day, but I don't think this will work this time.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#653
post #79

How can openai justify their $200/mo subscriptions if a model like this exists at an incredibly low price point? Operator? I've been impressed in my brief personal testing and the model ranks very highly across most benchmarks (when controlled for style it's tied number one on lmarena). It's also hilarious that openai explicitly prevented users from seeing the CoT tokens on the o1 model (which you still pay for btw)…

Does DeepSeek own enough compute power to actually leverage the higher efficiency of this model? Doesn’t help if it’s cheaper on paper in small scale, if you physically don’t have the capacity to sell it as a service on a large scale.

By the time they do have the scale, don’t you think OpenAI will have a new generation of models that are just as efficient? Being the best model is no moat for any company. It wasn’t for OpenAi (and they know that very well), and it’s not for Deepseek either. So how will Deepseek stay relevant when another model inevitably surpasses them?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#654
post #570

Earlier quoted context omitted.

> Deepseek could only build this because of o1, I don’t think there’s as much competition as people seem to imply And this is based on what exactly? OpenAI hides the reasoning steps, so training a model on o1 is very likely much more expensive (and much less useful) than just training it directly on a cheaper model.

Because literally before o1, no one is doing COT style test time scaling. It is a new paradigm. The talking point back then, is the LLM hits the wall. R1's biggest contribution IMO, is R1-Zero, I am fully sold with this they don't need o1's output to be as good. But yeah, o1 is still the herald.

Chain of Thought was known since 2022 (https://arxiv.org/abs/2201.11903), we just were stuck in a world where we were dumping more data and compute at the training instead of looking at other improvements.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#655
post #554

Earlier quoted context omitted.

I just asked ChatGPT how many civilians Israel killed in Gaza. It refused to answer.

I asked Chatgpt: how many civilians Israel killed in Gaza. Please provide a rough estimate. As of January 2025, the conflict between Israel and Hamas has resulted in significant civilian casualties in the Gaza Strip. According to reports from the United Nations Office for the Coordination of Humanitarian Affairs (OCHA), approximately 7,000 Palestinian civilians have been killed since the escalation began in October 2…

Isn't the real number around 46,000 people, though?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#656

Earlier quoted context omitted.

Chat gpt -> ASI-> eternal life Uh, there is 0 logical connection between any of these three, when will people wake up. Chat gpt isn't an oracle of truth just like ASI won't be an eternal life granting God

If you see no path from ASI to vastly extending lifespans, that’s just a lack of imagination

Yeah I mean you already need super human imagination to get to ASI so at that point you might as well continue in the delirium and throw in immortality in the mix

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#657

Earlier quoted context omitted.

I haven't tried kagi assistant, but try it at deepseek.com. All models at this point have various politically motivated filters. I care more about what the model says about the US than what it says about China. Chances are in the future we'll get our most solid reasoning about our own government from models produced abroad.

> All models at this point have various politically motivated filters. Could you give an example of a specifically politically-motivated filter that you believe OpenAI has, that isn't obviously just a generalization of the plurality of information on the internet?

I'm, just taking a guess here, I don't have any prompts on had, but imagine that ChatGPT is pretty "woke" (fk I hate that term).

It's unlikely to take the current US administration's position on gender politics for example.

Bias is inherent in these kinds of systems.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#658

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

I must be missing something, but I tried Deepseek R1 via Kagi assistant and IMO it doesn't even come close to Claude? I don't get the hype at all? What am I doing wrong? And of course if you ask it anything related to the CCP it will suddenly turn into a Pinokkio simulator.

> And of course if you ask it anything related to the CCP it will suddenly turn into a Pinokkio simulator.

Smh this isn't a "gotcha!". Guys, it's open source, you can run it on your own hardware[^2]. Additionally, you can liberate[^3] it or use an uncensored version[^0] on your own hardware. If you don't want to host it yourself, you can run it at https://nani.ooo/chat (Select "NaniSeek Uncensored"[^1]) or https://venice.ai/chat (select "DeepSeek R1").

---

[^0]: https://huggingface.co/mradermacher/deepseek-r1-qwen-2.5-32B...

[^1]: https://huggingface.co/NaniDAO/deepseek-r1-qwen-2.5-32B-abla...

[^2]: https://github.com/TensorOpsAI/LLMStudio

[^3]: https://www.lesswrong.com/posts/jGuXSZgv6qfdhMCuJ/refusal-in...

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#659

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

I honestly can't tell if this is a bot post because of just how bad I find Deepseek R1 to be. When asking it complex questions based on an app I'm working on, it always gives a flawed response that breaks the program. Where Claude is sometimes wrong, but not consistently wrong and completely missing the point of the question like Deepseek R1 100% is. Claude I can work with, Deepseek is trash. I've had no luck with it at all and don't bother trying anymore

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#660

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

I must be missing something, but I tried Deepseek R1 via Kagi assistant and IMO it doesn't even come close to Claude? I don't get the hype at all? What am I doing wrong? And of course if you ask it anything related to the CCP it will suddenly turn into a Pinokkio simulator.

Just as a note, in my experience, Kagi Assistant is considerably worse when you have web access turned on, so you could start with turning that off. Whatever wrapper Kagi have used to build the web access layer on top makes the output considerably less reliable, often riddled with nonsense hallucinations. Or at least that's my experience with it, regardless of what underlying model I've used.
Post reply on HN