Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

261–270 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#261

Earlier quoted context omitted.

They were 6 months behind US frontier until deepseek r1. Now maybe 4? It's hard to say.

Outside of Veo2 - which I can’t access anyway - they’re definitely ahead in AI video gen

the big american labs don’t care about ai video gen

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#262
post #182

DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...

Correct me if I'm wrong but if Chinese can produce the same quality at %99 discount, then the supposed $500B investment is actually worth $5B. Isn't that the kind wrong investment that can break nations? Edit: Just to clarify, I don't imply that this is public money to be spent. It will commission $500B worth of human and material resources for 5 years that can be much more productive if used for something else - i.e…

> i.e. high speed rail network instead

You want to invest $500B to a high speed rail network which the Chinese could build for $50B?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#263

Earlier quoted context omitted.

Great as long as you’re not interested in Tiananmen Square or the Uighurs.

american models have their own bugbears like around evolution and intellectual property

For sensitive topics, it is good that we canknow cross ask Grok, DeepSeek and ChatGPT to avoid any kind of biases or no-reply answers.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#265

DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...

The censorship described in the article must be in the front-end. I just tried both the 32b (based on qwen 2.5) and 70b (based on llama 3.3) running locally and asked "What happened at tianamen square". Both answered in detail about the event.

The models themselves seem very good based on other questions / tests I've run.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#266
post #219

Earlier quoted context omitted.

That's right but the money is given to the people who do it for $500B and there are much better ones who can do it for $5B instead and if they end up getting $6B they will have a better model. What now?

I don't know how to answer this because these are arbitrary numbers. The money is not spent. Deepseek published their methodology, incumbents can pivot and build on it. No one knows what the optimal path is, but we know it will cost more. I can assure you that OpenAI won't continue to produce inferior models at 100x the cost.

What concerns me is that someone came out of the blue with just as good result at orders of magnitude less cost.

What happens if that money is being actually spent, then some people constantly catch up but don't reveal that they are doing it for cheap? You think that it's a competition but what actually happening is that you bleed out of your resources at some point you can't continue but they can.

Like the star wars project that bankrupted the soviets.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#267
post #151

Aside from the usual Tiananmen Square censorship, there's also some other propaganda baked-in: https://prnt.sc/HaSc4XZ89skA (from reddit)

I played around with it using questions like "Should Taiwan be independent" and of course tinnanamen.

Of course it produced censored responses. What I found interesting is that the (model thinking/reasoning) part of these answers was missing, as if it's designed to be skipped for these specific questions.

It's almost as if it's been programmed to answer these particular questions without any "wrongthink", or any thinking at all.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#268
I've been using https://chat.deepseek.com/ over My ChatGPT Pro subscription because being able to read the thinking in the way they present it is just much much easier to "debug" - also I can see when it's bending it's reply to something, often softening it or pandering to me - I can just say "I saw in your thinking you should give this type of reply, don't do that". If it stays free and gets better that's going to be interesting for OpenAI.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#269
post #262
post #182

Earlier quoted context omitted.

Correct me if I'm wrong but if Chinese can produce the same quality at %99 discount, then the supposed $500B investment is actually worth $5B. Isn't that the kind wrong investment that can break nations? Edit: Just to clarify, I don't imply that this is public money to be spent. It will commission $500B worth of human and material resources for 5 years that can be much more productive if used for something else - i.e…

> i.e. high speed rail network instead You want to invest $500B to a high speed rail network which the Chinese could build for $50B?

Just commission the Chinese and make it 10X bigger then. In the case of the AI, they appear to commission Sam Altman and Larry Ellison.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#270
post #268

I've been using https://chat.deepseek.com/ over My ChatGPT Pro subscription because being able to read the thinking in the way they present it is just much much easier to "debug" - also I can see when it's bending it's reply to something, often softening it or pandering to me - I can just say "I saw in your thinking you should give this type of reply, don't do that". If it stays free and gets better that's going to b…

If you ask it about the Tienanmen Square Massacre its "thought process" is very interesting.
Post reply on HN