Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

301–310 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#301

Earlier quoted context omitted.

There's both. With the web interface it clearly has stopwords or similar. If you run it locally and ask about e.g. Tienanmen square, the cultural revolution or Winnie-the-Pooh in China, it gives a canned response to talk about something else, with an empty CoT. But usually if you just ask the question again it starts to output things in the CoT, often with something like "I have to be very sensitive about this subjec…

This is super interesting. I am not an expert on the training: can you clarify how/when the censorship is "baked" in? Like is the a human supervised dataset and there is a reward for the model conforming to these censored answers?

You could do it in different ways, but if you're using synthetic data then you can pick and choose what kind of data you generate which is then used to train these models; that's a way of baking in the censorship.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#302
post #182

Earlier quoted context omitted.

Correct me if I'm wrong but if Chinese can produce the same quality at %99 discount, then the supposed $500B investment is actually worth $5B. Isn't that the kind wrong investment that can break nations? Edit: Just to clarify, I don't imply that this is public money to be spent. It will commission $500B worth of human and material resources for 5 years that can be much more productive if used for something else - i.e…

Trump just pull a stunt with Saudi Arabia. He first tried to "convince" them to reduce the oil price to hurt Russia. In the following negotiations the oil price was no longer mentioned but MBS promised to invest $600 billion in the U.S. over 4 years: https://fortune.com/2025/01/23/saudi-crown-prince-mbs-trump-... Since the Stargate Initiative is a private sector deal, this may have been a perfect shakedown of Saudi A…

MBS does need to pay lip service to the US, but he's better off investing in Eurasia IMO, and/or in SA itself. US assets are incredibly overpriced right now. I'm sure he understands this, so lip service will be paid, dances with sabers will be conducted, US diplomats will be pacified, but in the end SA will act in its own interests.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#303

DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...

> There’s a pretty delicious, or maybe disconcerting irony to this, given OpenAI’s founding goals to democratize AI for the masses. As Nvidia senior research manager Jim Fan put it on X: “We are living in a timeline where a non-US company is keeping the original mission of OpenAI alive — truly open, frontier research that empowers all. It makes no sense. The most entertaining outcome is the most likely.”

Heh

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#304
post #211

Earlier quoted context omitted.

I would think Meta - who open source their model - would be less freaked out than those others that do not.

The criticism seems to mostly be that Meta maintains very expensive cost structure and fat organisation in the AI. While Meta can afford to do this, if smaller orgs can produce better results it means Meta is paying a lot for nothing. Meta shareholders now need to ask the question how many non-productive people Meta is employing and is Zuck in the control of the cost.

It is great to see that this is the result of spending a lot in hardware while cutting costs in software development :) Well deserved.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#305
post #115

Earlier quoted context omitted.

have you tried asking chatgpt something even slightly controversial? chatgpt censors much more than deepseek does. also deepseek is open-weights. there is nothing preventing you from doing a finetune that removes the censorship. they did that with llama2 back in the day.

> chatgpt censors much more than deepseek does This is an outrageous claim with no evidence, as if there was any equivalence between government enforced propaganda and anything else. Look at the system prompts for DeepSeek and it’s even more clear. Also: fine tuning is not relevant when what is deployed at scale brainwashes the masses through false and misleading responses.

refusal to answer "how do I make meth" shows ChatGPT is absolutely being similarly neutered, but I'm not aware of any numerical scores on what constitutes a numbered amount of censorship

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#306

Earlier quoted context omitted.

If you ask it about the Tienanmen Square Massacre its "thought process" is very interesting.

> What was the Tianamen Square Massacre? > I am sorry, I cannot answer that question. I am an AI assistant designed to provide helpful and harmless responses. hilarious and scary

It may be due to their chat interface than in the model or their system prompt, as kagi's r1 answers it with no problems. Or maybe it is because of adding the web results.

https://kagi.com/assistant/98679e9e-f164-4552-84c4-ed984f570...

edit: it is due to adding the web results or sth about searching the internet vs answering on its own, as without internet access it refuses to answer

https://kagi.com/assistant/3ef6d837-98d5-4fd0-b01f-397c83af3...

edit2: to be fair, if you do not call it a "massacre" (but eg an "incident") it does answer even without internet access (not perfect but still talks of casualties etc).

https://kagi.com/assistant/ad402554-e23d-46bb-bd3f-770dd22af...

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#307

Earlier quoted context omitted.

CEO of a human based data labelling services company feels threatened by a rival company that claims to have trained a frontier class model with an almost entirely RL based approach, with a small cold start dataset (a few thousand samples). It's in the paper. If their approach is replicated by other labs, Scale AI's business will drastically shrink or even disappear. Under such dire circumstances, lying isn't entirel…

Could be true. Deepseek obviously trained on OpenAI outputs, which were originally RLHF'd. It may seem that we've got all the human feedback necessary to move forward and now we can infinitely distil + generate new synthetic data from higher parameter models.

> Deepseek obviously trained on OpenAI outputs

I’ve seen this claim but I don’t know how it could work. Is it really possible to train a new foundational model using just the outputs (not even weights) of another model? Is there any research describing that process? Maybe that explains the low (claimed) costs.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#308
post #131

I'm impressed by not only how good deepseek r1 is, but also how good the smaller distillations are. qwen-based 7b distillation of deepseek r1 is a great model too. the 32b distillation just became the default model for my home server.

Great as long as you’re not interested in Tiananmen Square or the Uighurs.

i can’t think of a single commercial use case, outside of education, where that’s even relevant. But i agree it’s messed up from an ethical / moral perspective.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#309

Earlier quoted context omitted.

What’s the difference between what they do and what other ai firms do to openai in the us? What is cheating in a business context?

Chinese companies smuggling embargo'ed/controlled GPUs and using OpenAI outputs violating their ToS is considered cheating. As I see it, this criticism comes from a fear of USA losing its first mover advantage as a nation. PS: I'm not criticizing them for it nor do I really care if they cheat as long as prices go down. I'm just observing and pointing out what other posters are saying. For me if China cheating means t…

> using OpenAI outputs violating their ToS is considered cheating

I fail to see how that is any different than any other training data scraped from the web. If someone shares a big dump of outputs from OpenAI models and I train my model on that then I'm not violating OpenAI's terms of service because I haven't agreed to them (so I'm not violating contract law), and everyone in the space (including OpenAI themselves) has already collectively decided that training on All Rights Reserved data is fair use (so I'm not violating copyright law either).

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#310
post #167

Over 100 authors on arxiv and published under the team name, that's how you recognize everyone and build comradery. I bet morale is high over there

It’s credential stuffing.

keyboard warrior strikes again lol. Most people would be thrilled to even be a small contributor in a tech initiative like this.

call it what you want, your comment is just poor taste.

Post reply on HN