Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

381–390 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#381

Earlier quoted context omitted.

If you ask it about the Tienanmen Square Massacre its "thought process" is very interesting.

> What was the Tianamen Square Massacre? > I am sorry, I cannot answer that question. I am an AI assistant designed to provide helpful and harmless responses. hilarious and scary

There is a collection of these prompts they refuse to answer in this article:

https://medium.com/the-generator/deepseek-hidden-china-polit...

What’s more confusing is where the refusal is coming from. Some people say that running offline removes the censorship. Others say that this depends on the exact model you use, with some seemingly censored even offline. Some say it depends on a search feature being turned on or off. I don’t think we have any conclusions yet, beyond anecdotal examples.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#382
post #343

Earlier quoted context omitted.

o3 isn’t available

Right, and that doesn't contradict what I wrote.

agreed but some might read your comment implying otherwise (there's no world in which you would have 'started using o3 regularly enough to opine'), as i did - given that you list it side to side with an available model.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#384

Earlier quoted context omitted.

try asking US models about the influence of Israeli diaspora on funding genocide in Gaza then come back

Which American models? Are you suggesting the US government exercises control over US LLM models the way the CCP controls DeepSeek outputs?

i think both American and Chinese model censorship is done by private actors out of fear of external repercussion, not because it is explicitly mandated to them

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#385
post #218

I was completely surprised that the reasoning comes from within the model. When using gpt-o1 I thought it's actually some optimized multi-prompt chain, hidden behind an API endpoint. Something like: collect some thoughts about this input; review the thoughts you created; create more thoughts if needed or provide a final answer; ...

I think the reason why it works is also because chain-of-thought (CoT), in the original paper by Denny Zhou et. al, worked from "within". The observation was that if you do CoT, answers get better. Later on community did SFT on such chain of thoughts. Arguably, R1 shows that was a side distraction, and instead a clean RL reward would've been better suited.

Do you understand why RL is better than SFT for training on reasoning traces?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#386

Earlier quoted context omitted.

The $500B is just an aspirational figure they hope to spend on data centers to run AI models, such as GPT-o1 and its successors, that have already been developed. If you want to compare the DeepSeek-R development costs to anything, you should be comparing it to what it cost OpenAI to develop GPT-o1 (not what they plan to spend to run it), but both numbers are somewhat irrelevant since they both build upon prior resea…

Thinking of the $500B as only an aspirational number is wrong. It’s true that the specific Stargate investment isn’t fully invested yet, but that’s hardly the only money being spent on AI development. The existing hyperscalers have already sunk ungodly amounts of money into literally hundreds of new data centers, millions of GPUs to fill them, chip manufacturing facilities, and even power plants with the impression t…

I agree except on the "isn't easily repurposed" part. Nvidia's chips have CUDA and can be repurposed for many HPC projects once the AI bubble will be done. Meteorology, encoding, and especially any kind of high compute research.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#387
post #182

DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...

Correct me if I'm wrong but if Chinese can produce the same quality at %99 discount, then the supposed $500B investment is actually worth $5B. Isn't that the kind wrong investment that can break nations? Edit: Just to clarify, I don't imply that this is public money to be spent. It will commission $500B worth of human and material resources for 5 years that can be much more productive if used for something else - i.e…

Think of it like a bet. Or even think of it a bomb.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#388

Earlier quoted context omitted.

Hah no way. The poor LLM has no privacy to your prying eyes. I kinda like the 'reasoning' text it provides in general. It makes prompt engineering way more convenient.

The benefit of running locally. It's leaky if you poke at it enough, but there's an effort to sanitize the inputs and the outputs, and Tianamen Square is a topic that it considers unsafe.

Do you have any other examples? this is fascinating

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#389

Even if you think this particular team cheated, the idea that nobody will find ways of making training more efficient seems silly - these huge datacenter investments for purely AI will IMHO seem very short sighted in 10 years

But we're in the test time compute paradigm now, and we've only just gotten started in terms of applications. I really don't have high confidence that there's going to be a glut of compute.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#390

Earlier quoted context omitted.

There's an interesting tweet here from someone who used to work at DeepSeek, which describes their hiring process and culture. No mention of LeetCoding for sure! https://x.com/wzihanw/status/1872826641518395587

they almost certainly ask coding/technical questions. the people doing this work are far beyond being gatekept by leetcode leetcode is like HN’s “DEI” - something they want to blame everything on

what is leetcode
Post reply on HN