Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

761–770 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#761
post #262
post #182

Earlier quoted context omitted.

Correct me if I'm wrong but if Chinese can produce the same quality at %99 discount, then the supposed $500B investment is actually worth $5B. Isn't that the kind wrong investment that can break nations? Edit: Just to clarify, I don't imply that this is public money to be spent. It will commission $500B worth of human and material resources for 5 years that can be much more productive if used for something else - i.e…

> i.e. high speed rail network instead You want to invest $500B to a high speed rail network which the Chinese could build for $50B?

My understanding of the problems with high speed rail in the US is more fundamental than money.

The problem is loose vs strong property rights.

We don't have the political will in the US to use eminent domain like we did to build the interstates. High speed rail ultimately needs a straight path but if you can't make property acquisitions to build the straight rail path then this is all a non-starter in the US.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#762
post #547

Earlier quoted context omitted.

I don't disagree, but the important point is that Deepseek showed that it's not just about CapEx, which is what the US firms were/are lining up to battle with. In my opinion there is something qualitatively better about Deepseek in spite of its small size, even compared to o1-pro, that suggests a door has been opened. GPUs are needed to rapidly iterate on ideas, train, evaluate, etc., but Deepseek has shown us that w…

Back in the day there were a lot of things that appeared not to be about capex because the quality of the capital was improving so quickly. Computers became obsolete after a year or two. Then the major exponential trends finished running their course and computers stayed useful for longer. At that point, suddenly AWS popped up and it turned out computing was all about massive capital investments. AI will be similar.…

True but it is unknown how much of the capital will be used for training vs experimenting vs hosting vs talent.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#763

Earlier quoted context omitted.

They're using it via fireworks.ai, which is the 685B model. https://fireworks.ai/models/fireworks/deepseek-r1

How do you know which version it is? I didn't see anything in that link.

because they wouldn’t call it r1 otherwise unless they were unethical (like ollama is)

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#764

Has anyone done a benchmark on these reasoning models compared to simply prompting "non-reasoning" LLMs with massive chain of thought? For example, a go to test I've used (but will have to stop using soon) is: "Write some JS code to find the smallest four digit prime number whose digits are in strictly descending order" That prompt, on its own, usually leads to an incorrect response with non-reasoning models. They al…

Anecdotally, the reasoning is more effective than what I can get out of Claude with my "think()" tool/prompt. I did have trouble with R1 (and o1) with output formatting in some tool commands though (I have the models output a JSON array of commands with optional raw strings for some parameters) -- whereas Claude did not have this issue. In some cases it would not use the RAW format or would add extra backslashes when nesting JSON, which Claude managed okay and also listened when I asked for RAW output in that case.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#765

Earlier quoted context omitted.

It should be. I think AMD has left a lot on the table with respect to competing in the space (probably to the point of executive negligence) and the new US laws will help create several new Chinese competitors. NVIDIA probably has a bit of time left as the market leader, but it's really due mostly to luck.

> NVIDIA probably has a bit of time left as the market leader, but it's really due mostly to luck. Look, I think NVIDIA is overvalued and AI hype has poisoned markets/valuations quite a bit. But if I set that aside, I can't actually say NVIDIA is in the position they're in due to luck. Jensen has seemingly been executing against a cohesive vision for a very long time. And focused early on on the software side of the…

> I can't actually say NVIDIA is in the position they're in due to luck

They aren't, end of story.

Even though I'm not a scientist in the space, I studied at EPFL in 2013 and researchers in the ML space could write to Nvidia about their research with their university email and Nvidia would send top-tier hardware for free.

Nvidia has funded, invested and supported in the ML space when nobody was looking and it's only natural that the research labs ended up writing tools around its hardware.

I don't think their moat will hold forever, especially among big tech that has the resources to optimize around their use case but it's only natural they enjoy such a headstart.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#766

Earlier quoted context omitted.

I must be missing something, but I tried Deepseek R1 via Kagi assistant and IMO it doesn't even come close to Claude? I don't get the hype at all? What am I doing wrong? And of course if you ask it anything related to the CCP it will suddenly turn into a Pinokkio simulator.

Chinese models get a lot of hype online, they cheat on benchmarks by using benchmark data in training, they definitely train on other models outputs that forbid training and in normal use their performance seem way below OpenAI and Anthropic. The CCP set a goal and their AI engineer will do anything they can to reach it, but the end product doesn't look impressive enough.

cope, r1 is the best public model for my private benchmark tasks

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#767
post #504

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

The nVidia market price could also be questionable considering how much cheaper DS is to run.

The improved efficiency of steam engines in the past did not reduce coal consumption; instead, it enabled people to accomplish more work with the same resource.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#768

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

Given this comment, I tried it.

It's no where close to Claude, and it's also not better than OpenAI.

I'm so confused as to how people judge these things.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#769

Everyone is trying to say its better than the biggest closed models. It feels like it has parity, but its not the clear winner. But, its free and open and the quant models are insane. My anecdotal test is running models on a 2012 mac book pro using CPU inference and a tiny amount of RAM. The 1.5B model is still snappy, and answered the strawberry question on the first try with some minor prompt engineering (telling i…

you’re probably running it on ollama.

ollama is doing the pretty unethical thing of lying about whether you are running r1, most of the models they have labeled r1 are actually entirely different models

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#770

Earlier quoted context omitted.

How much memory do you have? I'm trying to figure out which is the best model to run on 48GB (unified memory).

32B works well (I have 48GB Macbook Pro M3)

you’re not running r1 dude.

e: no clue why i’m downvoted for this

Post reply on HN