Live data from Hacker News

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

arxiv.org

941–950 of 1001 posts

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#941

Earlier quoted context omitted.

CEO of a human based data labelling services company feels threatened by a rival company that claims to have trained a frontier class model with an almost entirely RL based approach, with a small cold start dataset (a few thousand samples). It's in the paper. If their approach is replicated by other labs, Scale AI's business will drastically shrink or even disappear. Under such dire circumstances, lying isn't entirel…

Could be true. Deepseek obviously trained on OpenAI outputs, which were originally RLHF'd. It may seem that we've got all the human feedback necessary to move forward and now we can infinitely distil + generate new synthetic data from higher parameter models.

Check the screenshot below re: training on OpenAI Outputs. They've fixed this since btw, but it's pretty obvious they used OpenAI outputs to train. I mean all the Open AI "mini" models are trained the same way. Hot take but feels like the AI labs are gonna gatekeep more models and outputs going forward.

https://x.com/ansonhw/status/1883510262608859181

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#942

Earlier quoted context omitted.

FWIW it works with Hide my Email, no issues there.

Thanks, but all the same I'm not going to jump through arbitrary hoops set up by people who think it's okay to just capriciously break email. They simply won't ever get me as a customer and/or advocate in the industry. Same thing goes for any business that is hostile toward open systems and standards.

Yup, I 100% get your point.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#943

For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…

[dead]

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#944

Earlier quoted context omitted.

Could you give an example of a prompt where this happened?

Here's one from yesterday. https://imgur.com/a/Dmoti0c Though I tried twice today and didn't get it again.

To be fair the "you" in that doesn't necessarily refer to either you or the AI.

It's a hypothetical which has the same answer for any object: human, AI, dog, flower.

You could more clearly write it as:

How many times would a person have to randomly change their name before they ended up with the name Claude?

The changes are totally random so it doesn't matter who is making them or what their original name was.

Try asking this instead:

If you start randomly changing each letter in your name, in order, to a another random letter, how many changes would it take before you ended up with the name "Claudeee"?

I added two extra e's to make the names the same length.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#945

The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.

CEO of Scale said Deepseek is lying and actually has a 50k GPU cluster. He said they lied in the paper because technically they aren't supposed to have them due to export laws. I feel like this is very likely. They obvious did some great breakthroughs, but I doubt they were able to train on so much less hardware.

R1 is double the size of o1. By that logic, shouldn’t o1 have been even cheaper to train?

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#946

Earlier quoted context omitted.

I love how people love throwing the word "left" as it means anything. Need I remind you how many times bots were caught on twitter using chatgpt praising putin? Sure, go ahead and call it left if it makes you feel better but I still take the European and American left over the left that is embedded into russia and china - been there, done that, nothing good ever comes out of it and deepseek is here to back me up with…

Some people feel reality has a leftwing bias.

Yes, people born after the fall of the USSR and the Berlin Wall, generally.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#947
post #302

Earlier quoted context omitted.

MBS does need to pay lip service to the US, but he's better off investing in Eurasia IMO, and/or in SA itself. US assets are incredibly overpriced right now. I'm sure he understands this, so lip service will be paid, dances with sabers will be conducted, US diplomats will be pacified, but in the end SA will act in its own interests.

One only needs to look as far back as the first Trump administration to see that Trump only cares about the announcement and doesn’t care about what’s actually done. And if you don’t want to look that far just lookup what his #1 donor Musk said…there is no actual $500Bn.

Yeah - Musk claims SoftBank "only" has $10B available for this atm.

There was an amusing interview with MSFT CEO Satya Nadella at Davos where he was asked about this, and his response was "I don't know, but I know I'm good for my $80B [that I'm investing to expand Azure]".

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#948

Earlier quoted context omitted.

I must be missing something, but I tried Deepseek R1 via Kagi assistant and IMO it doesn't even come close to Claude? I don't get the hype at all? What am I doing wrong? And of course if you ask it anything related to the CCP it will suddenly turn into a Pinokkio simulator.

Just as a note, in my experience, Kagi Assistant is considerably worse when you have web access turned on, so you could start with turning that off. Whatever wrapper Kagi have used to build the web access layer on top makes the output considerably less reliable, often riddled with nonsense hallucinations. Or at least that's my experience with it, regardless of what underlying model I've used.

That makes sense. When I used Kagi assistant 6 months ago I was able to jailbreak what it saw from the web results and it was given much less data from the actual web sites than Perplexity, just very brief excerpts to look at. I'm not overly impressed with Perplexity's web search capabilities either, but it was the better of the two.

Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

#949

Worth noting that people have been unpacking and analyzing DeepSeek-R1 vigorously for days already on X before it got to Hacker News — it wasn't always this way.

HN has a general tech audience including SWEs who are paid so much that they exhibit the Nobel Disease and fauxtrepeneurs who use AI as a buzzword. They exist on X too but the conversations are diffused. You’ll have a section of crypto bros on there who know nothing technical they are talking then. Other user’s algorithms will fit their level of deep technical familiarity with AI.
Post reply on HN