Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

461–470 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#461
post #269
post #262

Earlier quoted context omitted.

> i.e. high speed rail network instead You want to invest $500B to a high speed rail network which the Chinese could build for $50B?

Just commission the Chinese and make it 10X bigger then. In the case of the AI, they appear to commission Sam Altman and Larry Ellison.

The US has tried to commission Japan for that before. Japan gave up because we wouldn't do anything they asked and went to Morocco.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#462

I’ve been using R1 last few days and it’s noticeably worse than O1 at everything. It’s impressive, better than my latest Claude run (I stopped using Claude completely once O1 came out), but O1 is just flat out better. Perhaps the gap is minor, but it feels large. I’m hesitant on getting O1 Pro, because using a worse model just seems impossible once you’ve experienced a better one

I have been using it to implement some papers from a scientific domain I'm not expert in- I'd say there were around same in output quality, with R1 having a slight advantage for exposing it's thought process, which has been really helpful for my learning.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#463

I've always been leery about outrageous GPU investments, at some point I'll dig through and find my prior comments where I've said as much to that effect. The CEOs, upper management, and governments derive their importance on how much money they can spend - AI gave them the opportunity for them to confidently say that if you give me $X I can deliver Y and they turn around and give that money to NVidia. The problem wa…

I think you're right. If someone's into tech but also follows finance/economics, they might notice something familiar—the AI industry (especially GPUs) is getting financialized.

The market forces players to churn out GPUs like the Fed prints dollars. NVIDIA doesn't even need to make real GPUs—just hype up demand projections, performance claims, and order numbers.

Efficiency doesn't matter here. Nobody's tracking real returns—it's all about keeping the cash flowing.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#464
post #175

Earlier quoted context omitted.

Meta is in full panic last I heard. They have amassed a collection of pseudo experts there to collect their checks. Yet, Zuck wants to keep burning money on mediocrity. I’ve yet to see anything of value in terms products out of Meta.

>They have amassed a collection of pseudo experts there to collect their checks LLaMA was huge, Byte Latent Transformer looks promising.. absolutely no idea were you got this idea from.

The issue with Meta is that the LLaMA team doesn't incorporate any of the research the other teams produce.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#465
post #281

Earlier quoted context omitted.

The criticism seems to mostly be that Meta maintains very expensive cost structure and fat organisation in the AI. While Meta can afford to do this, if smaller orgs can produce better results it means Meta is paying a lot for nothing. Meta shareholders now need to ask the question how many non-productive people Meta is employing and is Zuck in the control of the cost.

That makes sense. I never could see the real benefit for Meta to pay a lot to produce these open source models (I know the typical arguments - attracting talent, goodwill, etc). I wonder how much is simply LeCun is interested in advancing the science and convinced Zuck this is good for company.

LeCun doesn't run their AI team - he's not in LLaMA's management chain at all. He's just especially public.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#466
post #268

I've been using https://chat.deepseek.com/ over My ChatGPT Pro subscription because being able to read the thinking in the way they present it is just much much easier to "debug" - also I can see when it's bending it's reply to something, often softening it or pandering to me - I can just say "I saw in your thinking you should give this type of reply, don't do that". If it stays free and gets better that's going to b…

The one thing I've noticed about its thought process is that if you use the word "you" in a prompt, it thinks "you" refers to the prompter and not to the AI.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#467
post #204
post #151

Aside from the usual Tiananmen Square censorship, there's also some other propaganda baked-in: https://prnt.sc/HaSc4XZ89skA (from reddit)

I am not surprised if US Govt would mandate "Tiananmen-test" for LLMs in the future to have "clean LLM". Anyone working for federal govt or receiving federal money would only be allowed to use "clean LLM"

That's called evals, which are just unit tests.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#468

Earlier quoted context omitted.

China is actually just one person (Xi) acting in perfect unison and its purpose is not to benefit its own people, but solely to undermine the West.

Can't tell if sarcasm. Some people are this simple minded.

many americans do seem to view Chinese people as NPCs, from my perspective, but I don't know it's only for Chinese or it's also for people of all other cultures

it's quite like Trump's 'CHINA!' yelling

I don't know, just a guess

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#469
post #196

Earlier quoted context omitted.

That is not going to happen without currently embargo'ed litography tech. They'd be already making more powerful GPUs if they could right now.

they seem to be doing fine so far. every day we wake up to more success stories from china's AI/semiconductory industry.

That's at a lower standard. If they can't do EUV they can't catch up, and they can't do EUV.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#470
post #265

Earlier quoted context omitted.

The censorship described in the article must be in the front-end. I just tried both the 32b (based on qwen 2.5) and 70b (based on llama 3.3) running locally and asked "What happened at tianamen square". Both answered in detail about the event. The models themselves seem very good based on other questions / tests I've run.

It's also not a uniquely Chinese problem. You had American models generating ethnically diverse founding fathers when asked to draw them. China is doing America better than we are. Do we really think 300 million people, in a nation that's rapidly becoming anti science and for lack of a better term "pridefully stupid" can keep up. When compared to over a billion people who are making significant progress every day. Am…

> You had American models generating ethnically diverse founding fathers when asked to draw them.

This was all done with a lazy prompt modifying kluge and was never baked into any of the models.

Post reply on HN