Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

241–250 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#241
post #219

Earlier quoted context omitted.

That's right but the money is given to the people who do it for $500B and there are much better ones who can do it for $5B instead and if they end up getting $6B they will have a better model. What now?

Are you under the impression it was some kind of fixed-scope contractor bid for a fixed price?

No, its just that those people intend to commission huge amount of people to build obscene amount of GPUs and put them together in an attempt to create a an unproven machine when others appear to be able to do it at the fraction of the cost.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#242

Earlier quoted context omitted.

Probably shouldn't be firing their blood boys just yet ... According to Musk, SoftBank only has $10B available for this atm.

Elon says a lot of things.

Funding secured!

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#243
post #151

Aside from the usual Tiananmen Square censorship, there's also some other propaganda baked-in: https://prnt.sc/HaSc4XZ89skA (from reddit)

Apparently the censorship isn't baked-in to the model itself, but rather is overlayed in the public chat interface. If you run it yourself, it is significantly less censored [0] [0] https://thezvi.substack.com/p/on-deepseeks-r1?open=false#%C2...

There's both. With the web interface it clearly has stopwords or similar. If you run it locally and ask about e.g. Tienanmen square, the cultural revolution or Winnie-the-Pooh in China, it gives a canned response to talk about something else, with an empty CoT. But usually if you just ask the question again it starts to output things in the CoT, often with something like "I have to be very sensitive about this subject" and "I have to abide by the guidelines", and typically not giving a real answer. With enough pushing it does start to converse about the issues somewhat even in the answers.

My guess is that it's heavily RLHF/SFT-censored for an initial question, but not for the CoT, or longer discussions, and the censorship has thus been "overfit" to the first answer.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#244
post #151

Aside from the usual Tiananmen Square censorship, there's also some other propaganda baked-in: https://prnt.sc/HaSc4XZ89skA (from reddit)

Just did a test with https://huggingface.co/bartowski/DeepSeek-R1-Distill-Qwen-32..., with the question "what happened at Tiananmen Square?", and here are parts of the thinking process:

  ...I also remember something about the "Tank Man" image, where a lone protester stood in front of a line of tanks. That image became iconic, symbolizing resistance against oppression. But I'm not sure what happened to that person or if they survived.

  After the crackdown, the government censored information about the event. So, within China, it's not openly discussed, and younger people might not know much about it because it's not taught in schools. But outside of China, it's a significant event in modern history, highlighting the conflict between authoritarian rule and the desire for democracy...

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#245
post #131

I'm impressed by not only how good deepseek r1 is, but also how good the smaller distillations are. qwen-based 7b distillation of deepseek r1 is a great model too. the 32b distillation just became the default model for my home server.

Great as long as you’re not interested in Tiananmen Square or the Uighurs.

american models have their own bugbears like around evolution and intellectual property

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#246
post #180

Earlier quoted context omitted.

Apparently the censorship isn't baked-in to the model itself, but rather is overlayed in the public chat interface. If you run it yourself, it is significantly less censored [0] [0] https://thezvi.substack.com/p/on-deepseeks-r1?open=false#%C2...

Oh, my experience was different. Got the model through ollama. I'm quite impressed how they managed to bake in the censorship. It's actually quite open about it. I guess censorship doesnt have as bad a rep in china as it has here? So it seems to me that's one of the main achievements of this model. Also another finger to anyone who said they can't publish their models cause of ethical reasons. Deepseek demonstrated c…

I mean US models are highly censored too.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#247
post #175

DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...

Meta is in full panic last I heard. They have amassed a collection of pseudo experts there to collect their checks. Yet, Zuck wants to keep burning money on mediocrity. I’ve yet to see anything of value in terms products out of Meta.

They got momentarily leap-frogged, which is how competition is supposed to work!

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#248
post #175

Earlier quoted context omitted.

Meta is in full panic last I heard. They have amassed a collection of pseudo experts there to collect their checks. Yet, Zuck wants to keep burning money on mediocrity. I’ve yet to see anything of value in terms products out of Meta.

I guess all that leetcoding and stack ranking didn't in fact produce "the cream of the crop"...

There's an interesting tweet here from someone who used to work at DeepSeek, which describes their hiring process and culture. No mention of LeetCoding for sure!

https://x.com/wzihanw/status/1872826641518395587

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#249
post #180

Earlier quoted context omitted.

Oh, my experience was different. Got the model through ollama. I'm quite impressed how they managed to bake in the censorship. It's actually quite open about it. I guess censorship doesnt have as bad a rep in china as it has here? So it seems to me that's one of the main achievements of this model. Also another finger to anyone who said they can't publish their models cause of ethical reasons. Deepseek demonstrated c…

> I guess censorship doesnt have as bad a rep in china as it has here It's probably disliked, just people know not to talk about it so blatantly due to chilling effects from aforementioned censorship. disclaimer: ignorant American, no clue what i'm talking about.

on the topic of censorship, US LLMs' censorship is called alignment. llama or ChatGPT's refusal on how to make meth or nuclear bombs is the same as not answering questions abput Tiananmen tank man as far as the matrix math word prediction box is concerned.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#250
post #146

Larry Ellison is 80. Masayoshi Son is 67. Both have said that anti-aging and eternal life is one of their main goals with investing toward ASI. For them it's worth it to use their own wealth and rally the industry to invest $500 billion in GPUs if that means they will get to ASI 5 years faster and ask the ASI to give them eternal life.

Larry especially has already invested in life-extension research.
Post reply on HN