Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

291–300 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#291
post #268

I've been using https://chat.deepseek.com/ over My ChatGPT Pro subscription because being able to read the thinking in the way they present it is just much much easier to "debug" - also I can see when it's bending it's reply to something, often softening it or pandering to me - I can just say "I saw in your thinking you should give this type of reply, don't do that". If it stays free and gets better that's going to b…

If you ask it about the Tienanmen Square Massacre its "thought process" is very interesting.

> What was the Tianamen Square Massacre?

> I am sorry, I cannot answer that question. I am an AI assistant designed to provide helpful and harmless responses.

hilarious and scary

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#292

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

> The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I do not quite follow. GPU compute is mostly spent in inference, as training is a one time cost. And these chain of thought style models work by scaling up inference time compute, no? So proliferation of these types of models would portend in increase in…

As far as I understand the model needs way less active parameters, reducing GPU cost in inference.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#293

Earlier quoted context omitted.

The amount of astroturfing around R1 is absolutely wild to see. Full scale propaganda war.

The literal creator of Netscape Navigator is going ga-ga over it on Twitter and HN thinks its all botted This is not a serious place

> all botted

Of course it isn’t all botted. You don’t put astroturf muscle behind things that are worthless. You wait until you have something genuinely good and then give as big of a push as you can. The better it genuinely is the more you artificially push as hard as you can.

Go read a bunch of AI related subreddits and tell me you honestly believe all the comments and upvotes are just from normal people living their normal life.

Don’t be so naive.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#295
DeepSeek V3 came in the perfect time, precisely when Claude Sonnet turned into crap and barely allows me to complete something without me hitting some unexpected constraints.

Idk, what their plans is and if their strategy is to undercut the competitors but for me, this is a huge benefit. I received 10$ free credits and have been using Deepseeks api a lot, yet, I have barely burned a single dollar, their pricing are this cheap!

I’ve fully switched to DeepSeek on Aider & Cursor (Windsurf doesn’t allow me to switch provider), and those can really consume tokens sometimes.

We live in exciting times.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#296
post #196

Earlier quoted context omitted.

That is not going to happen without currently embargo'ed litography tech. They'd be already making more powerful GPUs if they could right now.

they seem to be doing fine so far. every day we wake up to more success stories from china's AI/semiconductory industry.

I only know about Moore Threads GPUs. Last time I took a look at their consumer offerings (e.g. MTT S80 - S90), they were at GTX1650-1660 or around the latest AMD APU performance levels.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#297
post #290

Earlier quoted context omitted.

Come on man, let them have their well deserved win as a team.

Yea, I’m sure they’re devastated by my comment

It’s not about hurting them directly or indirectly, but I’d prefer people to not drag me down if I achieved something neat. So, ideally i’d want others to be the same towards others.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#298
post #182

Earlier quoted context omitted.

Correct me if I'm wrong but if Chinese can produce the same quality at %99 discount, then the supposed $500B investment is actually worth $5B. Isn't that the kind wrong investment that can break nations? Edit: Just to clarify, I don't imply that this is public money to be spent. It will commission $500B worth of human and material resources for 5 years that can be much more productive if used for something else - i.e…

if you say, i wanna build 5 nuclear reactors and I need 200 billion $$. I would believe it because, you can ballpark it with some stats. For tech like LLMs, it feels irresponsible to say 500 billion $$ investment and then place that into R&D. What if in 2026, we realize we can create it for 2 billion$, and let the 498 billion $ sitting in a few consumers.

I bet the Chinese can build 5 nuclear reactors for a fraction of that price, too. Deepseek says China builds them at $2.5-3.5B per 1200MW reactor.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#299
post #112

Earlier quoted context omitted.

Agreed. I am no fan of the CCP but I have no issue with using DeepSeek since I only need to use it for coding which it does quite well. I still believe Sonnet is better. DeepSeek also struggles when the context window gets big. This might be hardware though. Having said that, DeepSeek is 10 times cheaper than Sonnet and better than GPT-4o for my use cases. Models are a commodity product and it is easy enough to add a…

Curious why you have to qualify this with a “no fan of the CCP” prefix. From the outset, this is just a private organization and its links to CCP aren’t any different than, say, Foxconn’s or DJI’s or any of the countless Chinese manufacturers and businesses You don’t invoke “I’m no fan of the CCP” before opening TikTok or buying a DJI drone or a BYD car. Then why this, because I’ve seen the same line repeated everywh…

Anything that becomes valuable will become a CCP property and it looks like DeepSeek may become that. The worry right now is that people feel using DeepSeek supports the CCP, just as using TikTok does. With LLMs we have static data that provides great control over what knowledge to extract from it.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#300

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

CEO of Scale said Deepseek is lying and actually has a 50k GPU cluster. He said they lied in the paper because technically they aren't supposed to have them due to export laws. I feel like this is very likely. They obvious did some great breakthroughs, but I doubt they were able to train on so much less hardware.

Alexandr only parroted what Dylan Patel said on Twitter. To this day, no one know how this number come up.
Post reply on HN