DeepSeek R2 launch stalled as CEO balks at progress
11–20 of 186 posts
Re: DeepSeek R2 launch stalled as CEO balks at progress
#12"We had difficulties accessing OpenAI, our data provider." /s
To me that does seem like a reasonable speculation, though unproven.
Re: DeepSeek R2 launch stalled as CEO balks at progress
#13Earlier quoted context omitted.
but deepseek doesn't actually need to host inference right if they opensource it? I don't see why these companies even bother to host inference. deepseek doesn't need outreach (everyone knows about them) and the huge demand for sota will force western companies to host them anyway.
Releasing the model has paid off handsomely with name recognition and making a significant geopolitical and cultural statement. But will they keep releasing the weights or do an OpenAI and come up with a reason they can't release them anymore? At the end of the day, even if they release the weights, they probably want to make money and leverage the brand by hosting the model API and the consumer mobile app.
Re: DeepSeek R2 launch stalled as CEO balks at progress
#14"We had difficulties accessing OpenAI, our data provider." /s
Re: DeepSeek R2 launch stalled as CEO balks at progress
#15They just recently released the r1-0528 model which was a massive upgrade over the original R1 and is roughly on par with the current best proprietary western models. Let them take their time on R2.
With this combo, I have no reason to use Claude/Gemini for anything.
People don't realize how good the new Deepseek model is.
Re: DeepSeek R2 launch stalled as CEO balks at progress
#16Honestly, AI progress suffers because of these export restrictions. An open source model that can compete with Gemini Pro 2.5 and o3 is good for the world, and good for AI
Re: DeepSeek R2 launch stalled as CEO balks at progress
#17"We had difficulties accessing OpenAI, our data provider." /s
This, my guess is OpenAI wised up after r1 and put safeguards in place for o3 that it didn't have for o1, hence the delay.
DeepSeek-R1 0528 performs almost as well as o3 in AI quality benchmarks. So, either OpenAI didn't restrict access, DeepSeek wasn't using OpenAI's output, or using OpenAI's output doesn't have a material impact in DeepSeek's performance.
https://artificialanalysis.ai/?models=gpt-4-1%2Co4-mini%2Co3...
Re: DeepSeek R2 launch stalled as CEO balks at progress
#18Earlier quoted context omitted.
Releasing the model has paid off handsomely with name recognition and making a significant geopolitical and cultural statement. But will they keep releasing the weights or do an OpenAI and come up with a reason they can't release them anymore? At the end of the day, even if they release the weights, they probably want to make money and leverage the brand by hosting the model API and the consumer mobile app.
If they continue to release the weights + detailed reports what they did, I seriously don't understand why. I mean it's cool. I just don't understand why. It's such a cut throat environment where every little bit of moat counts. I don't think they're naive. I think I'm naive.
Now they are firmly on the map, which presumably helps with hiring, doing deals, influence. If they stop publishing something, they run the risk of being labelled a one-hit wonder who got lucky.
If they have a reason to believe they can do even better in the near future, releasing current tech might make sense.
Re: DeepSeek R2 launch stalled as CEO balks at progress
#19Honestly, AI progress suffers because of these export restrictions. An open source model that can compete with Gemini Pro 2.5 and o3 is good for the world, and good for AI
Your views on this question are going to differ a lot depending on the probability you assign to a conflict with China in the next five years. I feel like that number should be offered up for scrutiny before a discussion on the cost vs benefits of export controls even starts.
Re: DeepSeek R2 launch stalled as CEO balks at progress
#20They just recently released the r1-0528 model which was a massive upgrade over the original R1 and is roughly on par with the current best proprietary western models. Let them take their time on R2.
At this point the only models I use are o3/o3-pro and R1-0528. The OpenAI model is better at handling data and drawing inferences, whereas the DeepSeek model is better at handling text as a thing in itself -- i.e. for all writing and editing tasks. With this combo, I have no reason to use Claude/Gemini for anything. People don't realize how good the new Deepseek model is.