Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

981–990 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#981

Earlier quoted context omitted.

DeepSeek really is taking out OpenAI at the knees. It's shocking that the first direct peer competition to OpenAI is also doing it for an order of magnitude less as a side project.

I just tried DeepSeek for the first time and immediately canceled my OpenAI subscription. Seeing the chain of thought is now just mandatory for me after one prompt. That is absolutely incredible in terms of my own understanding of the question I asked. Even the chat UI feels better and less clunky. Now picture 20 years from now when the Chinese companies have access to digital Yuan transaction data along with all the…

I will probably sound like an idiot for saying this but I tested ChatGpt-o1 model against DeepSeek and came away not blown away. It seems like its comparable to OpenAI 4o but many here make it seems like it has eclipsed anything OpenAI has put out?

I asked it a simple question about the music from a 90s movie I liked as a child. Specifically to find the song that plays during a certain scene. The answer is a little tricky because in the official soundtrack the song is actually part of a larger arrangement and the song only starts playing X minutes into that specific track on the soundtrack album.

DeepSeek completely hallucinated a nonsense answer making up a song that didn't even exist in the movie or soundtrack and o1 got me more or less to the answer(it was 99% correct in that it got the right track but only somewhat close to the actual start time: it was off by 15 seconds).

Furthermore, the chain of thought of DeepSeek was impressive...in showing me how it it hallucinated but the chain of thought in o1 also led me to a pretty good thought process on how it derived the song I was looking for(and also taught me how a style of song called a "stinger" can be used to convey a sudden change in tone in the movie).

Maybe its like how Apple complains when users don't use their products right, im not using it right with these nonsense requests. :D

Both results tell me that DeepSeek needs more refinement and that OpenAI still cannot be trusted to fully replace a human because the answer still needed verification and correction despite being generally right.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#982

I wonder if Xai is sweating their imminent Grok 3 release because of DeepSeek. It’ll be interesting to see how good that model is.

Was Grok2 or Grok 1 any good? I thought Musk was a distant last place shipping garbage?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#983
post #146

Larry Ellison is 80. Masayoshi Son is 67. Both have said that anti-aging and eternal life is one of their main goals with investing toward ASI. For them it's worth it to use their own wealth and rally the industry to invest $500 billion in GPUs if that means they will get to ASI 5 years faster and ask the ASI to give them eternal life.

Mark my words, anything that comes of anti-aging will ultimately turn into a subscription to living.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#984
post #965
post #487

Earlier quoted context omitted.

I asked this > What was the Tianamen Square Event? The model went on a thinking parade about what happened (I couldn't read it all as it was fast) and as it finished its thinking, it removed the "thinking" and output > Sorry, I'm not sure how to approach this type of question yet. Let's chat about math, coding, and logic problems instead! Based on this, I'd guess the model is not censored but the platform is. Edit: r…

It's clearly trained to be a censor and an extension of the CCPs social engineering apparatus. Ready to be plugged into RedNote and keep the masses docile and focused on harmless topics.

Well. Let’s see how long ChstGPT will faithfully answer questions about Trump‘s attempted self-coup and the criminals that left nine people dead. Sometimes it’s better to be careful with the bold superiority.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#985

Earlier quoted context omitted.

The $500B is just an aspirational figure they hope to spend on data centers to run AI models, such as GPT-o1 and its successors, that have already been developed. If you want to compare the DeepSeek-R development costs to anything, you should be comparing it to what it cost OpenAI to develop GPT-o1 (not what they plan to spend to run it), but both numbers are somewhat irrelevant since they both build upon prior resea…

Thinking of the $500B as only an aspirational number is wrong. It’s true that the specific Stargate investment isn’t fully invested yet, but that’s hardly the only money being spent on AI development. The existing hyperscalers have already sunk ungodly amounts of money into literally hundreds of new data centers, millions of GPUs to fill them, chip manufacturing facilities, and even power plants with the impression t…

What is the rationale for "isn't easily repurposed"?

The hardware can train LLM but also be used for vision, digital twin, signal detection, autonomous agents, etc.

Military uses seem important too.

Can the large GPU based data centers not be repurposed to that?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#986
post #268

I've been using https://chat.deepseek.com/ over My ChatGPT Pro subscription because being able to read the thinking in the way they present it is just much much easier to "debug" - also I can see when it's bending it's reply to something, often softening it or pandering to me - I can just say "I saw in your thinking you should give this type of reply, don't do that". If it stays free and gets better that's going to b…

Chatgpt does this as well, it just doesn't display it in the UI. You can click on the "thinking" to expand and read the tomhought process.

No, ChatGPT o1 only shows you the summary. The real thought process is hidden. However, DeepSeek shows you the full thought process.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#987

Earlier quoted context omitted.

The chain of thought is super useful in so many ways, helping me: (1) learn, way beyond the final answer itself, (2) refine my prompt, whether factually or stylistically, (3) understand or determine my confidence in the answer.

useful but also annoying, I don't like the childish style of writing full of filler words etc.

It uses them as tokens to direct the chain of thought, and it is pretty interesting that it uses just those works specifically. Remember that this behavior was not hard-coded to the system.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#988
post #311
post #146

Larry Ellison is 80. Masayoshi Son is 67. Both have said that anti-aging and eternal life is one of their main goals with investing toward ASI. For them it's worth it to use their own wealth and rally the industry to invest $500 billion in GPUs if that means they will get to ASI 5 years faster and ask the ASI to give them eternal life.

Side note: I’ve read enough sci-fi to know that letting rich people live much longer than not rich is a recipe for a dystopian disaster. The world needs incompetent heirs to waste most of their inheritance, otherwise the civilization collapses to some kind of feudal nightmare.

What’s a good sci fi book about that?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#989

I've been comparing R1 to O1 and O1-pro, mostly in coding, refactoring and understanding of open source code. I can say that R1 is on par with O1. But not as deep and capable as O1-pro. R1 is also a lot more useful than Sonnete. I actually haven't used Sonnete in awhile. R1 is also comparable to the Gemini Flash Thinking 2.0 model, but in coding I feel like R1 gives me code that works without too much tweaking. I oft…

How do you pass these models code bases?

Some of the interfaces can realtime check websites

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#990

Earlier quoted context omitted.

Weird to see straight up Chinese propaganda on HN, but it’s a free platform in a free country I guess. Try posting an opposite dunking on China on a Chinese website.

Weird to see we've put out non stop anti Chinese propaganda for the last 60 years instead of addressing our issues here.

There are ignorant people everywhere. There are brilliant people everywhere.

Governments should be criticized when they do bad things. In America, you can talk openly about things you don’t like that the government has done. In China, you can’t. I know which one I’d rather live in.

Post reply on HN