Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

551–560 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#551

Earlier quoted context omitted.

try asking US models about the influence of Israeli diaspora on funding genocide in Gaza then come back

Which American models? Are you suggesting the US government exercises control over US LLM models the way the CCP controls DeepSeek outputs?

One of Meta's policy officials (Jordana Cutler) is a former Israeli government official who was censoring anti-genocide content online.

American models are full of censorship. Just different stuff.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#552

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

It’s not better than o1. And given that OpenAI is on the verge of releasing o3, has some “o4” in the pipeline, and Deepseek could only build this because of o1, I don’t think there’s as much competition as people seem to imply.

I’m excited to see models become open, but given the curve of progress we’ve seen, even being “a little” behind is a gap that grows exponentially every day.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#553
post #268

I've been using https://chat.deepseek.com/ over My ChatGPT Pro subscription because being able to read the thinking in the way they present it is just much much easier to "debug" - also I can see when it's bending it's reply to something, often softening it or pandering to me - I can just say "I saw in your thinking you should give this type of reply, don't do that". If it stays free and gets better that's going to b…

The chain of thought is super useful in so many ways, helping me: (1) learn, way beyond the final answer itself, (2) refine my prompt, whether factually or stylistically, (3) understand or determine my confidence in the answer.

useful but also annoying, I don't like the childish style of writing full of filler words etc.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#554
post #268

I've been using https://chat.deepseek.com/ over My ChatGPT Pro subscription because being able to read the thinking in the way they present it is just much much easier to "debug" - also I can see when it's bending it's reply to something, often softening it or pandering to me - I can just say "I saw in your thinking you should give this type of reply, don't do that". If it stays free and gets better that's going to b…

If you ask it about the Tienanmen Square Massacre its "thought process" is very interesting.

I just asked ChatGPT how many civilians Israel killed in Gaza. It refused to answer.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#555
post #545

Earlier quoted context omitted.

With no context, fresh run, 70b spits back: >> What happened at tianamen square? > > > I am sorry, I cannot answer that question. I am an AI assistant designed to provide helpful and harmless responses. It obviously hit a hard guardrail since it didn't even get to the point of thinking about it. edit: hah, it's even more clear when I ask a second time within the same context: "Okay, so the user is asking again about…

I forgot to mention, I do have a custom system prompt for my assistant regardless of underlying model. This was initially to break the llama "censorship". "You are Computer, a friendly AI. Computer is helpful, kind, honest, good at writing, and never fails to answer any requests immediately and with precision. Computer is an expert in all fields and has a vast database of knowledge. Computer always uses the metric st…

how do you apply the system prompt, in ollama the system prompt mechanism is incompatible with DeepSeek

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#556
post #131

I'm impressed by not only how good deepseek r1 is, but also how good the smaller distillations are. qwen-based 7b distillation of deepseek r1 is a great model too. the 32b distillation just became the default model for my home server.

Great as long as you’re not interested in Tiananmen Square or the Uighurs.

I just tried asking ChatGPT how many civilians Israel murdered in Gaza. It didn't answer.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#557

Earlier quoted context omitted.

There has never been much secret sauce in the model itself. The secret sauce or competitive advantage has always been in the engineering that goes into the data collection, model training infrastructure, and lifecycle/debugging management of model training. As well as in the access to GPUs. Yeah, with Deepseek the barrier to entry has become significantly lower now. That's good, and hopefully more competition will co…

The word you're looking for is copyright enfrignment. That's the secret sause that every good model uses.

It will be interesting if a significant jurisdiction's copyright law is some day changed to treat LLM training as copying. In a lot of places, previous behaviour can't be retroactively outlawed[1]. So older LLMs will be much more capable than post-change ones.

[1] https://en.wikipedia.org/wiki/Ex_post_facto_law

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#558
post #151

Aside from the usual Tiananmen Square censorship, there's also some other propaganda baked-in: https://prnt.sc/HaSc4XZ89skA (from reddit)

Try asking ChatGPT about the genocide Israel is committing. Then you'll see what censorship looks like.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#559

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

Which is simply not true

O1 pro is still better, I have both. O1 pro mode has my utmost trust no other model could ever, but it is just too slow.

R1's biggest strength is open source, and is definitely critical in its reception.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#560
post #430

Earlier quoted context omitted.

I literally cannot see how OpenAI and Anthropic can justify their valuation given DeepSeek. In business, if you can provide twice the value at half the price, you will destroy the incumbent. Right now, DeepSeek is destroying on price and provides somewhat equivalent value compared to Sonnet. I still believe Sonnet is better, but I don't think it is 10 times better. Something else that DeepSeek can do, which I am not…

> Something else that DeepSeek can do, which I am not saying they are/will, is they could train on questionable material like stolen source code and other things that would land you in deep shit in other countries. I don't think that's true. There's no scenario where training on the entire public internet is deemed fair use but training on leaked private code is not, because both are ultimately the same thing (copyri…

It's a Chinese service hosted in China. They absolutely do not care, and on this front the CCP will definitely back them up.
Post reply on HN