Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

891–900 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#891

Everyone is trying to say its better than the biggest closed models. It feels like it has parity, but its not the clear winner. But, its free and open and the quant models are insane. My anecdotal test is running models on a 2012 mac book pro using CPU inference and a tiny amount of RAM. The 1.5B model is still snappy, and answered the strawberry question on the first try with some minor prompt engineering (telling i…

you’re probably running it on ollama. ollama is doing the pretty unethical thing of lying about whether you are running r1, most of the models they have labeled r1 are actually entirely different models

If you’re referring to what I think you’re referring to, those distilled models are from deepseek and not ollama https://github.com/deepseek-ai/DeepSeek-R1

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#892

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

I tried two questions that I had recently asked o1 pro mode. The first was about setting up a GitHub action to build a Hugo website. I provided it with the config code, and asked it about setting the directory to build from. It messed this up big time and decided that I should actually be checking out the git repo to that directory instead. I can see in the thinking section that it’s actually thought of the right sol…

I’ve had the exact opposite experience. But mine was in using both models to propose and ultimately write a refactor. If you don’t get this type of thing on the first shot with o1 pro you’re better off opening up a new chat, refining your prompt, and trying again. Soon as your asks get smaller within this much larger context I find it gets lost and starts being inconsistent in its answers. Even when the task remains the same as the initial prompt it starts coming up with newer more novel solutions halfway through implementation.

R1 seems much more up to the task of handling its large context window and remaining consistent. The search experience is also a lot better than search capable OpenAI models. It doesn’t get as stuck in a search response template and can answer questions in consideration of it.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#893
post #605

Earlier quoted context omitted.

False equivalency. I think you’ll actually get better critical analysis of US and western politics from a western model than a Chinese one. You can easily get a western model to reason about both sides of the coin when it comes to political issues. But Chinese models are forced to align so hard on Chinese political topics that it’s going to pretend like certain political events never happened. E.g try getting them to…

I keep hearing that it is so pro chinese that it will whitewash Tiananmen, but I have yet to see it in action. Here it is on both of the topics you asked about. AFAICT, it is pretty fair views on both. R1 14b quantized running locally on Tiananmen Square: Alright, the user is asking for more detailed information about the 1989 Tiananmen Square protests and what's referred to as a "massacre." From our previous convers…

Firstly, "R1 14b quantized"? You mean a quantised DeepSeek-R1-Distill-Qwen-14B? That is Qwen 2.5, it is not DeepSeek v3. Surely they didn't finetune Qwen to add more censorship.

Secondly, most of the censorship is a filter added on top of the model when run through chat.deepseek.com (and I've no idea about system prompt), it is only partially due to the actual model's training data.

Also, I'd rather people didn't paste huge blocks of text into HN comments.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#894

Earlier quoted context omitted.

Whenever I use it, it just seems to spin itself in circles for ages, spit out a half-assed summary and give up. Is it like the OpenAI models in that in needs to be prompted in extremely-specific ways to get it to not be garbage?

O1 doesn’t seem to need any particularly specific prompts. It seems to work just fine on just about anything I give it. It’s still not fantastic, but often times it comes up with things I either would have had to spend a lot of time to get right or just plainly things I didn’t know about myself.

I don’t ask LLMs about anything going on in my personal or business life. It’s purely a technical means to an end for me. So that’s where the disconnect is maybe.

For what I’m doing OpenAI’s models consistently rank last. I’m even using Flash 2 over 4o mini.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#895

Earlier quoted context omitted.

I don't find this to be true at all, maybe it has a few niche advantages, but GPT has significantly more data (which is what people are using these things for), and honestly, if GPT-5 comes out in the next month or two, people are likely going to forget about deepseek for a while. Also, I am incredibly suspicious of bot marketing for Deepseek, as many AI related things have. "Deepseek KILLED ChatGPT!", "Deepseek just…

the unpleasant truth is that the odious "bot marketing" you perceive is just the effect of influencers everywhere seizing upon the exciting topic du jour if you go back a few weeks or months there was also hype about minimax, nvidia's "world models", dsv3, o3, hunyuan, flux, papers like those for titans or lcm rendering transformers completely irrelevant… the fact that it makes for better "content" than usual (say fo…

Thanks for saying it. People are far too cynical, and blame everything on bots. The truth is they should be a lot more cynical, and blame everything on human tendencies!

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#896
post #605

Earlier quoted context omitted.

False equivalency. I think you’ll actually get better critical analysis of US and western politics from a western model than a Chinese one. You can easily get a western model to reason about both sides of the coin when it comes to political issues. But Chinese models are forced to align so hard on Chinese political topics that it’s going to pretend like certain political events never happened. E.g try getting them to…

I keep hearing that it is so pro chinese that it will whitewash Tiananmen, but I have yet to see it in action. Here it is on both of the topics you asked about. AFAICT, it is pretty fair views on both. R1 14b quantized running locally on Tiananmen Square: Alright, the user is asking for more detailed information about the 1989 Tiananmen Square protests and what's referred to as a "massacre." From our previous convers…

14b isn't the model being discussed here.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#897

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

So long as you don't ask it about tiananmen square 1989. Or Tibet. Or Taiwan. Or the Xinjiang internment camps. Just a few off the top of my head but thousands of others if you decide to dive deep. You get a shrug at best. Which does beg the question what responses you'd get in certain contexts.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#898

Earlier quoted context omitted.

Assuming you are US citizen, you should be worried about USG, not CCP. CCP having your data could rarely hurt you, unlike your own government. So gemini, chatgpt and so are more dangerous for you in a way.

Central EU citizen. I don't know, I am not naive about US and privacy, but as far as I know, US's motivation is mostly profit, not growth at absolutely any (human) cost, human rights repression, and world dominance.

[dead]

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#899

Earlier quoted context omitted.

Most people I talked with don't grasp how big of an event this is. I consider is almost as similar to as what early version of linux did to OS ecosystem.

I never tried the $200 a month subscription but it just solved a problem for me that neither o1 or claude was able to solve and did it for free. I like everything about it better. All I can think is "Wait, this is completely insane!"

Something off about this comment and the account it belongs to being 7 days old. Please post the problem/prompt you used so it can be cross checked.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#900

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

So long as you don't ask it about tiananmen square 1989. Or Tibet. Or Taiwan. Or the Xinjiang internment camps. Just a few off the top of my head but thousands of others if you decide to dive deep. You get a shrug at best. Which does beg the question what responses you'd get in certain contexts.

EDIT: I was incorrect, this does not work on the 14b model (and I presume above)

Works fine locally. Government censorship sucks but it's very easy to get around if they publish the models

Post reply on HN