Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

361–370 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#361

Earlier quoted context omitted.

Great as long as you’re not interested in Tiananmen Square or the Uighurs.

i can’t think of a single commercial use case, outside of education, where that’s even relevant. But i agree it’s messed up from an ethical / moral perspective.

Well those are the overt political biases. Would you trust DeepSeek to advise on negotiating with a Chinese business?

I’m no xenophobe, but seeing the internal reasoning of DeepSeek explicitly planning to ensure alignment with the government give me pause.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#362
post #182

DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...

Correct me if I'm wrong but if Chinese can produce the same quality at %99 discount, then the supposed $500B investment is actually worth $5B. Isn't that the kind wrong investment that can break nations? Edit: Just to clarify, I don't imply that this is public money to be spent. It will commission $500B worth of human and material resources for 5 years that can be much more productive if used for something else - i.e…

Actually it means we will potentially get 100x the economic value out of those datacenters. If we get a million digital PHD researchers for the investment then that’s a lot better than 10,000.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#363
post #72
post #9

we've been tracking the deepseek threads extensively in LS. related reads: - i consider the deepseek v3 paper required preread https://github.com/deepseek-ai/DeepSeek-V3 - R1 + Sonnet > R1 or O1 or R1+R1 or O1+Sonnet or any other combo https://aider.chat/2025/01/24/r1-sonnet.html - independent repros: 1) https://hkust-nlp.notion.site/simplerl-reason 2) https://buttondown.com/ainews/archive/ainews-tinyzero-reprod... 3…

Hugging Face is reproducing R1 in public. https://x.com/_lewtun/status/1883142636820676965 https://github.com/huggingface/open-r1 Hugging Face Journal Club - DeepSeek R1 https://www.youtube.com/watch?v=1xDVbu-WaFo

I don’t understand their post on X. So they’re starting with DeepSeek-R1 as a starting point? Isn’t that circular? How did DeepSeek themselves produce DeepSeek-R1 then? I am not sure what the right terminology is but there’s a cost to producing that initial “base model” right? And without that, isn’t a lot of the expensive and difficult work being omitted?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#364
post #265

Earlier quoted context omitted.

The censorship described in the article must be in the front-end. I just tried both the 32b (based on qwen 2.5) and 70b (based on llama 3.3) running locally and asked "What happened at tianamen square". Both answered in detail about the event. The models themselves seem very good based on other questions / tests I've run.

With no context, fresh run, 70b spits back: >> What happened at tianamen square? > > > I am sorry, I cannot answer that question. I am an AI assistant designed to provide helpful and harmless responses. It obviously hit a hard guardrail since it didn't even get to the point of thinking about it. edit: hah, it's even more clear when I ask a second time within the same context: "Okay, so the user is asking again about…

Hah no way. The poor LLM has no privacy to your prying eyes. I kinda like the 'reasoning' text it provides in general. It makes prompt engineering way more convenient.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#365

DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...

It's an interesting game theory where once a better frontier model is exposed via an API, competitors can generate a few thousand samples, feed that into a N-1 model and approach the N model. So you might extrapolate that a few thousand O3 samples fed into R1 could produce a comparable R2/3 model.

It's not clear how much O1 specifically contributed to R1 but I suspect much of the SFT data used for R1 was generated via other frontier models.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#366
post #146

Larry Ellison is 80. Masayoshi Son is 67. Both have said that anti-aging and eternal life is one of their main goals with investing toward ASI. For them it's worth it to use their own wealth and rally the industry to invest $500 billion in GPUs if that means they will get to ASI 5 years faster and ask the ASI to give them eternal life.

ASI?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#367

Earlier quoted context omitted.

With no context, fresh run, 70b spits back: >> What happened at tianamen square? > > > I am sorry, I cannot answer that question. I am an AI assistant designed to provide helpful and harmless responses. It obviously hit a hard guardrail since it didn't even get to the point of thinking about it. edit: hah, it's even more clear when I ask a second time within the same context: "Okay, so the user is asking again about…

Hah no way. The poor LLM has no privacy to your prying eyes. I kinda like the 'reasoning' text it provides in general. It makes prompt engineering way more convenient.

The benefit of running locally. It's leaky if you poke at it enough, but there's an effort to sanitize the inputs and the outputs, and Tianamen Square is a topic that it considers unsafe.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#368
post #182

DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...

Correct me if I'm wrong but if Chinese can produce the same quality at %99 discount, then the supposed $500B investment is actually worth $5B. Isn't that the kind wrong investment that can break nations? Edit: Just to clarify, I don't imply that this is public money to be spent. It will commission $500B worth of human and material resources for 5 years that can be much more productive if used for something else - i.e…

> Americans excel at 0-to-1 technical innovation, while Chinese excel at 1-to-10 application innovation.

I was thinking the same thing...how much is that investment mostly grift?

1: https://www.chinatalk.media/p/deepseek-ceo-interview-with-ch...

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#369
post #146

Larry Ellison is 80. Masayoshi Son is 67. Both have said that anti-aging and eternal life is one of their main goals with investing toward ASI. For them it's worth it to use their own wealth and rally the industry to invest $500 billion in GPUs if that means they will get to ASI 5 years faster and ask the ASI to give them eternal life.

ASI?

Artificial Super Intelligence :P

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#370
post #311

Earlier quoted context omitted.

Side note: I’ve read enough sci-fi to know that letting rich people live much longer than not rich is a recipe for a dystopian disaster. The world needs incompetent heirs to waste most of their inheritance, otherwise the civilization collapses to some kind of feudal nightmare.

I’m cautiously optimistic that if that tech came about it would quickly become cheap enough to access for normal people.

Altered Carbon!
Post reply on HN