Earlier quoted context omitted.
CEO of a human based data labelling services company feels threatened by a rival company that claims to have trained a frontier class model with an almost entirely RL based approach, with a small cold start dataset (a few thousand samples). It's in the paper. If their approach is replicated by other labs, Scale AI's business will drastically shrink or even disappear. Under such dire circumstances, lying isn't entirel…
Could be true. Deepseek obviously trained on OpenAI outputs, which were originally RLHF'd. It may seem that we've got all the human feedback necessary to move forward and now we can infinitely distil + generate new synthetic data from higher parameter models.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
941–950 of 1001 posts
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#942Earlier quoted context omitted.
FWIW it works with Hide my Email, no issues there.
Thanks, but all the same I'm not going to jump through arbitrary hoops set up by people who think it's okay to just capriciously break email. They simply won't ever get me as a customer and/or advocate in the industry. Same thing goes for any business that is hostile toward open systems and standards.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#943For those who haven't realized it yet, Deepseek-R1 is better than claude 3.5 and better than OpenAI o1-pro, better than Gemini. It is simply smarter -- a lot less stupid, more careful, more astute, more aware, more meta-aware, etc. We know that Anthropic and OpenAI and Meta are panicking. They should be. The bar is a lot higher now. The justification for keeping the sauce secret just seems a lot more absurd. None of…
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#944Earlier quoted context omitted.
Could you give an example of a prompt where this happened?
Here's one from yesterday. https://imgur.com/a/Dmoti0c Though I tried twice today and didn't get it again.
It's a hypothetical which has the same answer for any object: human, AI, dog, flower.
You could more clearly write it as:
How many times would a person have to randomly change their name before they ended up with the name Claude?
The changes are totally random so it doesn't matter who is making them or what their original name was.
Try asking this instead:
If you start randomly changing each letter in your name, in order, to a another random letter, how many changes would it take before you ended up with the name "Claudeee"?
I added two extra e's to make the names the same length.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#945The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.
CEO of Scale said Deepseek is lying and actually has a 50k GPU cluster. He said they lied in the paper because technically they aren't supposed to have them due to export laws. I feel like this is very likely. They obvious did some great breakthroughs, but I doubt they were able to train on so much less hardware.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#946Earlier quoted context omitted.
I love how people love throwing the word "left" as it means anything. Need I remind you how many times bots were caught on twitter using chatgpt praising putin? Sure, go ahead and call it left if it makes you feel better but I still take the European and American left over the left that is embedded into russia and china - been there, done that, nothing good ever comes out of it and deepseek is here to back me up with…
Some people feel reality has a leftwing bias.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#947Earlier quoted context omitted.
MBS does need to pay lip service to the US, but he's better off investing in Eurasia IMO, and/or in SA itself. US assets are incredibly overpriced right now. I'm sure he understands this, so lip service will be paid, dances with sabers will be conducted, US diplomats will be pacified, but in the end SA will act in its own interests.
One only needs to look as far back as the first Trump administration to see that Trump only cares about the announcement and doesn’t care about what’s actually done. And if you don’t want to look that far just lookup what his #1 donor Musk said…there is no actual $500Bn.
There was an amusing interview with MSFT CEO Satya Nadella at Davos where he was asked about this, and his response was "I don't know, but I know I'm good for my $80B [that I'm investing to expand Azure]".
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#948Earlier quoted context omitted.
I must be missing something, but I tried Deepseek R1 via Kagi assistant and IMO it doesn't even come close to Claude? I don't get the hype at all? What am I doing wrong? And of course if you ask it anything related to the CCP it will suddenly turn into a Pinokkio simulator.
Just as a note, in my experience, Kagi Assistant is considerably worse when you have web access turned on, so you could start with turning that off. Whatever wrapper Kagi have used to build the web access layer on top makes the output considerably less reliable, often riddled with nonsense hallucinations. Or at least that's my experience with it, regardless of what underlying model I've used.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#949Worth noting that people have been unpacking and analyzing DeepSeek-R1 vigorously for days already on X before it got to Hacker News — it wasn't always this way.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#950OpenAI is bust and will go bankrupt. The red flags have been there the whole time. Now it is just glaringly obvious. The AI bubble has burst!!!