Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

321–330 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#321

Earlier quoted context omitted.

> I guess censorship doesnt have as bad a rep in china as it has here It's probably disliked, just people know not to talk about it so blatantly due to chilling effects from aforementioned censorship. disclaimer: ignorant American, no clue what i'm talking about.

My guess would be that most Chinese even support the censorship at least to an extent for its stabilizing effect etc. CCP has quite a high approval rating in China even when it's polled more confidentially. https://dornsife.usc.edu/news/stories/chinese-communist-part...

Yep. And invent a new type of VPN every quarter to break free.

The indifferent mass prevails in every country, similarly cold to the First Amendment and Censorship. And engineers just do what they love to do, coping with reality. Activism is not for everyone.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#322
post #265

DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...

The censorship described in the article must be in the front-end. I just tried both the 32b (based on qwen 2.5) and 70b (based on llama 3.3) running locally and asked "What happened at tianamen square". Both answered in detail about the event. The models themselves seem very good based on other questions / tests I've run.

I think the guardrails are just very poor. If you ask it a few times with clear context, the responses are mixed.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#323
post #131

I'm impressed by not only how good deepseek r1 is, but also how good the smaller distillations are. qwen-based 7b distillation of deepseek r1 is a great model too. the 32b distillation just became the default model for my home server.

Great as long as you’re not interested in Tiananmen Square or the Uighurs.

try asking US models about the influence of Israeli diaspora on funding genocide in Gaza then come back

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#324

Earlier quoted context omitted.

> The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I do not quite follow. GPU compute is mostly spent in inference, as training is a one time cost. And these chain of thought style models work by scaling up inference time compute, no? So proliferation of these types of models would portend in increase in…

As far as I understand the model needs way less active parameters, reducing GPU cost in inference.

If you don't need so many gpu calcs regardless of how you get there, maybe nvidia loses money from less demand (or stock price), or there are more wasted power companies in the middle of no where (extremely likely), and maybe these dozen doofus almost trillion dollar ai companies also out on a few 100 billion of spending.

So it's not the end of the world. Look at the efficiency of databases from the mid 1970s to now. We have figured out so many optimizations and efficiencies and better compression and so forth. We are just figuring out what parts of these systems are needed.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#325

I've always been leery about outrageous GPU investments, at some point I'll dig through and find my prior comments where I've said as much to that effect. The CEOs, upper management, and governments derive their importance on how much money they can spend - AI gave them the opportunity for them to confidently say that if you give me $X I can deliver Y and they turn around and give that money to NVidia. The problem wa…

[deleted]

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#326

Earlier quoted context omitted.

I guess all that leetcoding and stack ranking didn't in fact produce "the cream of the crop"...

There's an interesting tweet here from someone who used to work at DeepSeek, which describes their hiring process and culture. No mention of LeetCoding for sure! https://x.com/wzihanw/status/1872826641518395587

Deepseek team is mostly quants from my understanding which explains why they were able to pull this off. Some of the best coders I’ve met have been quants.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#327
post #165

Earlier quoted context omitted.

In Communist theoretical texts the term "propaganda" is not negative and Communists are encouraged to produce propaganda to keep up morale in their own ranks and to produce propaganda that demoralize opponents. The recent wave of the average Chinese has a better quality of life than the average Westerner propaganda is an obvious example of propaganda aimed at opponents.

Is it propaganda if it's true?

I haven't been to China since 2019, but it is pretty obvious that median quality of life is higher in the US. In China, as soon as you get out of Beijing-Shanghai-Guangdong cities you start seeing deep poverty, people in tiny apartments that are falling apart, eating meals in restaurants that are falling apart, and the truly poor are emaciated. Rural quality of life is much higher in the US.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#328
post #146

Larry Ellison is 80. Masayoshi Son is 67. Both have said that anti-aging and eternal life is one of their main goals with investing toward ASI. For them it's worth it to use their own wealth and rally the industry to invest $500 billion in GPUs if that means they will get to ASI 5 years faster and ask the ASI to give them eternal life.

Can we wait until our political systems aren't putting 80+ year olds in charge BEFORE we cure aging?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#329
post #175

DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...

Meta is in full panic last I heard. They have amassed a collection of pseudo experts there to collect their checks. Yet, Zuck wants to keep burning money on mediocrity. I’ve yet to see anything of value in terms products out of Meta.

What I don't understand is why Meta needs so many VPs and directors. Shouldn't the model R&D be organized holacratically? The key is to experiment as many ideas as possible anyway. Those who can't experiment or code should remain minimal in such a fast-pacing area.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#330

Earlier quoted context omitted.

What’s the difference between what they do and what other ai firms do to openai in the us? What is cheating in a business context?

Chinese companies smuggling embargo'ed/controlled GPUs and using OpenAI outputs violating their ToS is considered cheating. As I see it, this criticism comes from a fear of USA losing its first mover advantage as a nation. PS: I'm not criticizing them for it nor do I really care if they cheat as long as prices go down. I'm just observing and pointing out what other posters are saying. For me if China cheating means t…

I understand that that’s what others are saying, but I think it’s very silly. We’re talking about international businesses, not kids on a playground. The rules are what you can get away with (same way openai can train on the open internet without anyone doing a thing).
Post reply on HN