Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

701–710 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#701
post #268

I've been using https://chat.deepseek.com/ over My ChatGPT Pro subscription because being able to read the thinking in the way they present it is just much much easier to "debug" - also I can see when it's bending it's reply to something, often softening it or pandering to me - I can just say "I saw in your thinking you should give this type of reply, don't do that". If it stays free and gets better that's going to b…

I tried signing up, but it gave me some bullshit "this email domain isn't supported in your region." I guess they insist on a GMail account or something? Regardless I don't even trust US-based LLM products to protect my privacy, let alone China-based. Remember kids: If it's free, you're the product. I'll give it a while longer before I can run something competitive on my own hardware. I don't mind giving it a few yea…

FWIW it works with Hide my Email, no issues there.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#702

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

Most people I talked with don't grasp how big of an event this is. I consider is almost as similar to as what early version of linux did to OS ecosystem.

Agreed: Worked on a tough problem in philosophy last night with DeepSeek on which I have previously worked with Claude. DeepSeek was at least as good and I found the output format better. I also did not need to provide a “pre-prompt” as I do with Claude.

And free use and FOSS.

Yep, game changer that opens the floodgates.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#705
post #572

Earlier quoted context omitted.

> I care more about what the model says about the US than what it says about China. This I don't get. If you want to use an LLM to take some of the work off your hands, I get it. But to ask an LLM for a political opinion?

I guess it matters if you're trying to build bots destined to your home country... More seriously, it doesn't have to be about political opinion. Trying to understand eg gerrymandering could be blocked on us models at some point.

Gerrymandering can simply be looked up in a dictionary or on wikipedia. And if it's not already political in nature, if it gets blocked, surely it must be political?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#706

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

While I agree its real competition are we so certain that R1 is indeed better? The times I have used it, its impressive but I would not throw it a title of the best model.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#707

OpenAI is bust and will go bankrupt. The red flags have been there the whole time. Now it is just glaringly obvious. The AI bubble has burst!!!

They just got 500 billion and they'll probably make that back in military contracts so this is unlikely (unfortunately)

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#708

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

While I agree its real competition are we so certain that R1 is indeed better? The times I have used it, its impressive but I would not throw it a title of the best model.

I'm sure it's not better in every possible way but after using it extensively over the weekend it seems a bit better than o1-pro, which was my previous pick for the top spot. The best part is that it catches itself going down an erroneous path and self-corrects.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#709

Earlier quoted context omitted.

This is really poor test though, of course the most recently trained model knows the newest libraries or knows that a library was renamed. Not disputing it's best at reasoning but you need a different test for that.

"recently trained" can't be an argument: those tools have to work with "current" data, otherwise they are useless.

That's a different part of the implementation details. If you were to break the system into mocroservices, the model is a binary blob with a mocroservices wrapper and accessing web search is another microservice entirely. You really don't want the entire web to be constantly compressed and re-released as a new model iteration, it's super inefficient.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#710
post #504

Earlier quoted context omitted.

The nVidia market price could also be questionable considering how much cheaper DS is to run.

It should be. I think AMD has left a lot on the table with respect to competing in the space (probably to the point of executive negligence) and the new US laws will help create several new Chinese competitors. NVIDIA probably has a bit of time left as the market leader, but it's really due mostly to luck.

> NVIDIA probably has a bit of time left as the market leader, but it's really due mostly to luck.

Look, I think NVIDIA is overvalued and AI hype has poisoned markets/valuations quite a bit. But if I set that aside, I can't actually say NVIDIA is in the position they're in due to luck.

Jensen has seemingly been executing against a cohesive vision for a very long time. And focused early on on the software side of the business to make actually using the GPUs easier. The only luck is that LLMs became popular.. but I would say consistent execution at NVIDIA is why they are the most used solution today.

Post reply on HN