How can openai justify their $200/mo subscriptions if a model like this exists at an incredibly low price point? Operator? I've been impressed in my brief personal testing and the model ranks very highly across most benchmarks (when controlled for style it's tied number one on lmarena). It's also hilarious that openai explicitly prevented users from seeing the CoT tokens on the o1 model (which you still pay for btw)…
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
251–260 of 1001 posts
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#252Over 100 authors on arxiv and published under the team name, that's how you recognize everyone and build comradery. I bet morale is high over there
Same thing happened to Google Gemini paper (1000+ authors) and it was described as big co promo culture (everyone wants credits). Interesting how narratives shift https://arxiv.org/abs/2403.05530
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#253Question about the rule-based rewards (correctness and format) mentioned in the paper: Does the raw base model just expected “stumble upon“ a correct answer /correct format to get a reward and start the learning process? Are there any more details about the reward modelling?
Good question. When BF Skinner used to train his pigeons, he’d initially reinforce any tiny movement that at least went in the right direction. For the exact reasons you mentioned. For example, instead of waiting for the pigeon to peck the lever directly (which it might not do for many hours), he’d give reinforcement if the pigeon so much as turned its head towards the lever. Over time, he’d raise the bar. Until, eve…
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#254Earlier quoted context omitted.
I would argue there is too little hype given the downloadable models for Deep Seek. There should be alot of hype around this organically. If anything, the other half good fully closed non ChatGPT models are astroturfing. I made a post in december 2023 whining about the non hype for Deep Seek. https://news.ycombinator.com/item?id=38505986
Possible for that to also be true! There’s a lot of astroturfing from a lot of different parties for a few different reasons. Which is all very interesting.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#255Perhaps the gap is minor, but it feels large. I’m hesitant on getting O1 Pro, because using a worse model just seems impossible once you’ve experienced a better one
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#256Question about the rule-based rewards (correctness and format) mentioned in the paper: Does the raw base model just expected “stumble upon“ a correct answer /correct format to get a reward and start the learning process? Are there any more details about the reward modelling?
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#257Aside from the usual Tiananmen Square censorship, there's also some other propaganda baked-in: https://prnt.sc/HaSc4XZ89skA (from reddit)
In Communist theoretical texts the term "propaganda" is not negative and Communists are encouraged to produce propaganda to keep up morale in their own ranks and to produce propaganda that demoralize opponents. The recent wave of the average Chinese has a better quality of life than the average Westerner propaganda is an obvious example of propaganda aimed at opponents.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#258Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#259The US Economy is pretty vulnerable here. If it turns out that you, in fact, don't need a gazillion GPUs to build SOTA models it destroys a lot of perceived value. I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.
Re: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
#260DeepSeek-R1 has apparently caused quite a shock wave in SV ... https://venturebeat.com/ai/why-everyone-in-ai-is-freaking-ou...
Correct me if I'm wrong but if Chinese can produce the same quality at %99 discount, then the supposed $500B investment is actually worth $5B. Isn't that the kind wrong investment that can break nations? Edit: Just to clarify, I don't imply that this is public money to be spent. It will commission $500B worth of human and material resources for 5 years that can be much more productive if used for something else - i.e…
https://fortune.com/2025/01/23/saudi-crown-prince-mbs-trump-...
Since the Stargate Initiative is a private sector deal, this may have been a perfect shakedown of Saudi Arabia. SA has always been irrationally attracted to "AI", so perhaps it was easy. I mean that part of the $600 billion will go to "AI".