Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

831–840 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#831

Earlier quoted context omitted.

I must be missing something, but I tried Deepseek R1 via Kagi assistant and IMO it doesn't even come close to Claude? I don't get the hype at all? What am I doing wrong? And of course if you ask it anything related to the CCP it will suddenly turn into a Pinokkio simulator.

I tried Deepseek R1 via Kagi assistant and it was much better than claude or gpt. I asked for suggestions for rust libraries for a certain task and the suggestions from Deepseek were better. Results here: https://x.com/larrysalibra/status/1883016984021090796

That's interesting!

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#832

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

I must be missing something, but I tried Deepseek R1 via Kagi assistant and IMO it doesn't even come close to Claude? I don't get the hype at all? What am I doing wrong? And of course if you ask it anything related to the CCP it will suddenly turn into a Pinokkio simulator.

They censor different things. Try asking any model from the west to write an erotic story and it will refuse. Deekseek has no trouble doing so.

Different cultures allow different things.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#833
Deepseek seems to create enormously long reasoning traces. I gave it the following for fun. It thought for a very long time (307 seconds), displaying a very long and stuttering trace before, losing confidence on the second part of the problem and getting it way wrong. GPTo1 got similarly tied in knots and took 193 seconds, getting the right order of magnitude for part 2 (0.001 inches). Gemini 2.0 Exp was much faster (it does not provide its reasoning time, but it was well under 60 second), with a linear reasoning trace, and answered both parts correctly.

I have a large, flat square that measures one mile on its side (so that it's one square mile in area). I want to place this big, flat square on the surface of the earth, with its center tangent to the surface of the earth. I have two questions about the result of this: 1. How high off the ground will the corners of the flat square be? 2. How far will a corner of the flat square be displaced laterally from the position of the corresponding corner of a one-square-mile area whose center coincides with the center of the flat area but that conforms to the surface of the earth?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#834
post #828

Earlier quoted context omitted.

I should’ve maybe been more explicit, it’s Claudes service that I think sucks atm, not their model. It feels like the free quota has been lowered much more than previously, and I have been using it since it was available to EU. I can’t count how many times I’ve started a conversation and after a couple of messages I get ”unexpected constrain (yada yada)”. It is either that or I get a notification saying ”defaulting t…

> Anthropic have hit their maximum capacity Yeah. They won't reset my API limit until February even though I have 50 dollars in funds that they can take from me. It looks like I may need to look at using Amazon instead.

> They won't reset my API limit until February even though I have 50 dollars in funds that they can take from me

That’s scummy.

I’ve heard good stuff about poe.com, have you looked at them?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#835

Earlier quoted context omitted.

How likely is this? Just a cursory probing of deepseek yields all kinds of censoring of topics. Isn't it just as likely Chinese sponsors of this have incentivized and sponsored an undercutting of prices so that a more favorable LLM is preferred on the market? Think about it, this is something they are willing to do with other industries. And, if LLMs are going to be engineering accelerators as the world believes, the…

I trust China a lot more than Meta and my own early tests do indeed show that Deepseek is far less censored than Llama.

Did you try asking deepseek about June 4th, 1989? Edit: it seems that basically the whole month of July 1989 is blocked. Any other massacres and genocides the model is happy to discuss.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#836

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

I don't find this to be true at all, maybe it has a few niche advantages, but GPT has significantly more data (which is what people are using these things for), and honestly, if GPT-5 comes out in the next month or two, people are likely going to forget about deepseek for a while. Also, I am incredibly suspicious of bot marketing for Deepseek, as many AI related things have. "Deepseek KILLED ChatGPT!", "Deepseek just…

the unpleasant truth is that the odious "bot marketing" you perceive is just the effect of influencers everywhere seizing upon the exciting topic du jour

if you go back a few weeks or months there was also hype about minimax, nvidia's "world models", dsv3, o3, hunyuan, flux, papers like those for titans or lcm rendering transformers completely irrelevant…

the fact that it makes for better "content" than usual (say for titans) is because of the competitive / political / "human interest" context — china vs the US, open weights vs not, little to no lip service paid to "safety" and "alignment" vs those being primary aspects of messaging and media strategy, export controls and allegedly low hardware resources vs tons of resources, election-related changes in how SV carries itself politically — and while that is to blame for the difference in sheer scale the underlying phenomenon is not at all different

the disease here is influencerism and the pus that oozes out of the sores it produces is rarely very organic

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#837
post #180

Earlier quoted context omitted.

Apparently the censorship isn't baked-in to the model itself, but rather is overlayed in the public chat interface. If you run it yourself, it is significantly less censored [0] [0] https://thezvi.substack.com/p/on-deepseeks-r1?open=false#%C2...

Oh, my experience was different. Got the model through ollama. I'm quite impressed how they managed to bake in the censorship. It's actually quite open about it. I guess censorship doesnt have as bad a rep in china as it has here? So it seems to me that's one of the main achievements of this model. Also another finger to anyone who said they can't publish their models cause of ethical reasons. Deepseek demonstrated c…

don't confuse the actual R1 (671b params) with the distilled models (the ones that are plausible to run locally.) Just as you shouldn't conclude about how o1 behaves when you are using o1-mini. maybe you're running the 671b model via ollama, but most folks here are not

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#838

Earlier quoted context omitted.

The one thing I've noticed about its thought process is that if you use the word "you" in a prompt, it thinks "you" refers to the prompter and not to the AI.

Could you give an example of a prompt where this happened?

Here's one from yesterday.

https://imgur.com/a/Dmoti0c

Though I tried twice today and didn't get it again.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#839

Earlier quoted context omitted.

aren't the smaller param models all just Qwen/Llama trained on R1 600bn?

yes, this is all ollamas fault

ollama is stating there's a difference: https://ollama.com/library/deepseek-r1

"including six dense models distilled from DeepSeek-R1 based on Llama and Qwen. "

people just don't read? not sure there's reason to criticize ollama here.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#840
post #605

Earlier quoted context omitted.

False equivalency. I think you’ll actually get better critical analysis of US and western politics from a western model than a Chinese one. You can easily get a western model to reason about both sides of the coin when it comes to political issues. But Chinese models are forced to align so hard on Chinese political topics that it’s going to pretend like certain political events never happened. E.g try getting them to…

GPT4 is also full of ideology, but of course the type you probably grew up with, so harder to see. (No offense intended, this is just the way ideology works). Try for example to persuade GPT to argue that the workers doing data labeling in Kenya should be better compensated relative to the programmers in SF, as the work they do is both critical for good data for training and often very gruesome, with many workers get…

[dead]
Post reply on HN