Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

871–880 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#871

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

I don't have access to o1-pro, but in my testing R1 performs noticably worse than o1.

It's more fun to use though because you can read the reasoning tokens live so I end up using it anyway.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#873

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

not sure why people are surprised, it's been known a long time that RLHF essentially lobotomizes LLMs by training them to give answers the base model wouldn't give. Deepseek is better because they didn't gimp their own model

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#874
post #219

Earlier quoted context omitted.

$500 billion is $500 billion. If new technology means we can get more for a dollar spent, then $500 billion gets more, not less.

That's right but the money is given to the people who do it for $500B and there are much better ones who can do it for $5B instead and if they end up getting $6B they will have a better model. What now?

The $500B wasnt given to the founders, investors and execs to do it better. It was given to them to enrich the tech exec and investor class. That's why it was that expensive - because of the middlemen who take enormous gobs of cash for themselves as profit and make everything more expensive. Precisely the same reason why everything in the US is more expensive.

Then the Open Source world came out of the left and b*tch slapped all those head honchos and now its like this.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#875
post #862

Earlier quoted context omitted.

> For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. Worse at writing. Its prose is overwrought. It's yet to learn that "less is more"

That's not what I've seen. See https://eqbench.com/results/creative-writing-v2/deepseek-ai_... , where someone fed it a large number of prompts. Weirdly, while the first paragraph from the first story was barely GPT-3 grade, 99% of the rest of the output blew me away (and is continuing to do so, as I haven't finished reading it yet.) I tried feeding a couple of the prompts to gpt-4o, o1-pro and the current Gemini 2.0…

What you linked is actually not good prose.

Good writing is how people speak.

Your example is overstuffed with similes.

Just because you can doesn't mean you should.

> He sauntered toward her

"sauntered" - nobody actually talks like this. Stuff like that on each paragraph.

It's fanficcy

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#876

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

I honestly can't tell if this is a bot post because of just how bad I find Deepseek R1 to be. When asking it complex questions based on an app I'm working on, it always gives a flawed response that breaks the program. Where Claude is sometimes wrong, but not consistently wrong and completely missing the point of the question like Deepseek R1 100% is. Claude I can work with, Deepseek is trash. I've had no luck with it…

It has a 64k context window. O1 has 128k Claude has 200k or 500K

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#877
post #554

Earlier quoted context omitted.

If you ask it about the Tienanmen Square Massacre its "thought process" is very interesting.

I just asked ChatGPT how many civilians Israel killed in Gaza. It refused to answer.

Why lie? I have asked ChatGPT some Gaza questions several times and it's actually surprisingly critical of Israel and the US.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#878
post #9

we've been tracking the deepseek threads extensively in LS. related reads: - i consider the deepseek v3 paper required preread https://github.com/deepseek-ai/DeepSeek-V3 - R1 + Sonnet > R1 or O1 or R1+R1 or O1+Sonnet or any other combo https://aider.chat/2025/01/24/r1-sonnet.html - independent repros: 1) https://hkust-nlp.notion.site/simplerl-reason 2) https://buttondown.com/ainews/archive/ainews-tinyzero-reprod... 3…

The discord invite link ( https://discord.gg/xJJMRaWCRt ) in ( https://www.latent.space/p/community ) is invalid

I had the same issue. Was able to use it to join via the discord app ("add a server").

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#879

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

> We know that Anthropic and OpenAI and Meta are panicking

Right after Altman turned OpenAI to private to boot...

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#880

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

I don't find this to be true at all, maybe it has a few niche advantages, but GPT has significantly more data (which is what people are using these things for), and honestly, if GPT-5 comes out in the next month or two, people are likely going to forget about deepseek for a while. Also, I am incredibly suspicious of bot marketing for Deepseek, as many AI related things have. "Deepseek KILLED ChatGPT!", "Deepseek just…

I think it's less bot marketing but more that a lot people hate C-suites. And a lot people hate the USA.

The narrative is the USA can never win. Even the whole AI trend was entirely started by the US companies, the moment a Chinese company publishes something resembling the SOTA it becomes the evidence of the fall of the USA.

Post reply on HN