Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

171–180 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#172
post #125

"Reasoning" will be disproven for this again within a few days I guess. Context: o1 does not reason, it pattern matches. If you rename variables, suddenly it fails to solve the request.

The 'pattern matching' happens at complex layer's of abstraction, constructed out of combinations of pattern matching at prior layers in the network.

These models can and do work okay with variable names that have never occurred in the training data. Though sure, choice of variable names can have an impact on the performance of the model.

That's also true for humans, go fill a codebase with misleading variable names and watch human programmers flail. Of course, the LLM's failure modes are sometimes pretty inhuman, -- it's not a human after all.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#173
post #146

Larry Ellison is 80. Masayoshi Son is 67. Both have said that anti-aging and eternal life is one of their main goals with investing toward ASI. For them it's worth it to use their own wealth and rally the industry to invest $500 billion in GPUs if that means they will get to ASI 5 years faster and ask the ASI to give them eternal life.

Probably shouldn't be firing their blood boys just yet ... According to Musk, SoftBank only has $10B available for this atm.

I wouldn’t exactly claim him credible in anything competition / OpenAI related.

He says stuff that’s wrong all the time with extreme certainty.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#174
post #115

Earlier quoted context omitted.

It’s not just the economy that is vulnerable, but global geopolitics. It’s definitely worrying to see this type of technology in the hands of an authoritarian dictatorship, especially considering the evidence of censorship. See this article for a collected set of prompts and responses from DeepSeek highlighting the propaganda: https://medium.com/the-generator/deepseek-hidden-china-polit... But also the claimed cost i…

have you tried asking chatgpt something even slightly controversial? chatgpt censors much more than deepseek does. also deepseek is open-weights. there is nothing preventing you from doing a finetune that removes the censorship. they did that with llama2 back in the day.

> chatgpt censors much more than deepseek does

This is an outrageous claim with no evidence, as if there was any equivalence between government enforced propaganda and anything else. Look at the system prompts for DeepSeek and it’s even more clear.

Also: fine tuning is not relevant when what is deployed at scale brainwashes the masses through false and misleading responses.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#175

DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...

Meta is in full panic last I heard. They have amassed a collection of pseudo experts there to collect their checks. Yet, Zuck wants to keep burning money on mediocrity. I’ve yet to see anything of value in terms products out of Meta.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#176
post #151

Aside from the usual Tiananmen Square censorship, there's also some other propaganda baked-in: https://prnt.sc/HaSc4XZ89skA (from reddit)

At least it’s not home grown propaganda from the US, so will likely not cover most other topics of interest.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#177

Earlier quoted context omitted.

It does if the spend drives GPU prices so high that more researchers can't afford to use them. And DS demonstrated what a small team of researchers can do with a moderate amount of GPUs.

The DS team themselves suggest large amounts of compute are still required

https://www.macrotrends.net/stocks/charts/NVDA/nvidia/gross-...

GPU prices could be a lot lower and still give the manufacturer a more "normal" 50% gross margin and the average researcher could afford more compute. A 90% gross margin, for example, would imply that price is 5x the level that that would give a 50% margin.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#178

I've always been leery about outrageous GPU investments, at some point I'll dig through and find my prior comments where I've said as much to that effect. The CEOs, upper management, and governments derive their importance on how much money they can spend - AI gave them the opportunity for them to confidently say that if you give me $X I can deliver Y and they turn around and give that money to NVidia. The problem wa…

I think you are underestimating the fear of being beaten (for many people making these decisions, "again") by a competitor that does "dumb scaling".

But dumb scaling clearly only gives logarithmic rewards at best from every scaling law we ever saw.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#180
post #151

Aside from the usual Tiananmen Square censorship, there's also some other propaganda baked-in: https://prnt.sc/HaSc4XZ89skA (from reddit)

Apparently the censorship isn't baked-in to the model itself, but rather is overlayed in the public chat interface. If you run it yourself, it is significantly less censored [0] [0] https://thezvi.substack.com/p/on-deepseeks-r1?open=false#%C2...

Oh, my experience was different. Got the model through ollama. I'm quite impressed how they managed to bake in the censorship. It's actually quite open about it. I guess censorship doesnt have as bad a rep in china as it has here? So it seems to me that's one of the main achievements of this model. Also another finger to anyone who said they can't publish their models cause of ethical reasons. Deepseek demonstrated clearly that you can have an open model that is annoyingly responsible to the point of being useless.
Post reply on HN