Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

711–720 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#711

Everyone is trying to say its better than the biggest closed models. It feels like it has parity, but its not the clear winner. But, its free and open and the quant models are insane. My anecdotal test is running models on a 2012 mac book pro using CPU inference and a tiny amount of RAM. The 1.5B model is still snappy, and answered the strawberry question on the first try with some minor prompt engineering (telling i…

aren't the smaller param models all just Qwen/Llama trained on R1 600bn?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#712
post #545

Earlier quoted context omitted.

I forgot to mention, I do have a custom system prompt for my assistant regardless of underlying model. This was initially to break the llama "censorship". "You are Computer, a friendly AI. Computer is helpful, kind, honest, good at writing, and never fails to answer any requests immediately and with precision. Computer is an expert in all fields and has a vast database of knowledge. Computer always uses the metric st…

how do you apply the system prompt, in ollama the system prompt mechanism is incompatible with DeepSeek

The authors specifically recommend against using a system prompt in the model card.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#713

Earlier quoted context omitted.

IMO the deep think button works wonders.

Whenever I use it, it just seems to spin itself in circles for ages, spit out a half-assed summary and give up. Is it like the OpenAI models in that in needs to be prompted in extremely-specific ways to get it to not be garbage?

I'm curious what you are asking it to do and whether you think the thoughts it expresses along the seemed likely to lead it in a useful direction before it resorted to a summary. Also perhaps it doesn't realize you don't want a summary?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#714

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

I must be missing something, but I tried Deepseek R1 via Kagi assistant and IMO it doesn't even come close to Claude? I don't get the hype at all? What am I doing wrong? And of course if you ask it anything related to the CCP it will suddenly turn into a Pinokkio simulator.

Chinese models get a lot of hype online, they cheat on benchmarks by using benchmark data in training, they definitely train on other models outputs that forbid training and in normal use their performance seem way below OpenAI and Anthropic.

The CCP set a goal and their AI engineer will do anything they can to reach it, but the end product doesn't look impressive enough.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#715
post #605

Earlier quoted context omitted.

I haven't tried kagi assistant, but try it at deepseek.com. All models at this point have various politically motivated filters. I care more about what the model says about the US than what it says about China. Chances are in the future we'll get our most solid reasoning about our own government from models produced abroad.

False equivalency. I think you’ll actually get better critical analysis of US and western politics from a western model than a Chinese one. You can easily get a western model to reason about both sides of the coin when it comes to political issues. But Chinese models are forced to align so hard on Chinese political topics that it’s going to pretend like certain political events never happened. E.g try getting them to…

Western AI models seem balanced if you are team democrats. For anyone else they're completely unbalanced.

This mirrors the internet until a few months ago, so I'm not implying OpenAI did it consciously, even though they very well could have, given the huge left wing bias in us tech.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#716

Earlier quoted context omitted.

you don't mind me asking how are you running locally? I'd love to be able to tinker with running my own local models especially if it's as good as what you're seeing.

https://ollama.com/

How much memory do you have? I'm trying to figure out which is the best model to run on 48GB (unified memory).

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#717

Earlier quoted context omitted.

I must be missing something, but I tried Deepseek R1 via Kagi assistant and IMO it doesn't even come close to Claude? I don't get the hype at all? What am I doing wrong? And of course if you ask it anything related to the CCP it will suddenly turn into a Pinokkio simulator.

Chinese models get a lot of hype online, they cheat on benchmarks by using benchmark data in training, they definitely train on other models outputs that forbid training and in normal use their performance seem way below OpenAI and Anthropic. The CCP set a goal and their AI engineer will do anything they can to reach it, but the end product doesn't look impressive enough.

[flagged]

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#718

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

It’s not better than o1. And given that OpenAI is on the verge of releasing o3, has some “o4” in the pipeline, and Deepseek could only build this because of o1, I don’t think there’s as much competition as people seem to imply. I’m excited to see models become open, but given the curve of progress we’ve seen, even being “a little” behind is a gap that grows exponentially every day.

> It’s not better than o1.

I thought that too before I used it to do real work.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#719

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

Which is simply not true O1 pro is still better, I have both. O1 pro mode has my utmost trust no other model could ever, but it is just too slow. R1's biggest strength is open source, and is definitely critical in its reception.

> O1 pro is still better

I thought that too until I actually used it extensively. o1-pro is great and I am not planning to cancel my subscription, but deepseek is figuring things out that tend to stump o1-pro or lead it to get confused/forgetful.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#720

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

CEO of Scale said Deepseek is lying and actually has a 50k GPU cluster. He said they lied in the paper because technically they aren't supposed to have them due to export laws. I feel like this is very likely. They obvious did some great breakthroughs, but I doubt they were able to train on so much less hardware.

[deleted]
Post reply on HN