[flagged]
OpenAI says it has evidence DeepSeek used its model to train competitor
871–880 of 1001 posts
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#872[flagged]
[flagged]
If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#873Earlier quoted context omitted.
Even if all that about training is true, the bigger cost is inference and Deepseek is 100x cheaper. That destroys OpenAI/Anthropic's value proposition of having a unique secret sauce so users are quickly fleeing to cheaper alternatives. Google Deepmind's recent Gemini 2.0 Flash Thinking is also priced at the new Deepseek level. It's pretty good (unlike previous Gemini models). [0] https://x.com/deedydas/status/188335…
I mean, Deepseek is currently charging 100x less. That doesn't tell us much about how cheaper it is to run inference on.
What is definitely true is that there are already other providers offering DeepSeek R1 (e.g. on OpenRouter[1]) for $7/m-in and $7/m-out. Meanwhile OpenAI is charging $15/m-in and $60/m-out. So already you're seeing at least 5x cheaper inference with R1 vs O1 with a bunch of confounding factors. But it is hard to say anything truly concrete about efficiency OpenAI does not disclose the actual compute required to run inference for O1.
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#874Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#875Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#876Let’s race to the bottom.
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#877Reading this post, I can’t help but wonder if people realize the irony in what they’re saying. 1. “The issue is when you [take it out of the platform and] are doing it to create your own model for your own purposes,” 2. “There’s a technique in AI called distillation . . . when one model learns from another model [and] kind of sucks the knowledge out of the parent model,”
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#878Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#879Called it from day 0, impossible to reach that performance with 5M, they had to distill OpenAI (or some other leading foundational model). Got downvoted to oblivion by people who haven't been told what to think by MSM yet. Now it's on FT and everywhere, good, what matters is that truth comes out eventually. I don't take any sides and think what DeepSeek did is fair play, however, what I do find harmful about this is,…
The "evidence" is very weak though: >The San Francisco-based ChatGPT maker told the Financial Times it had seen some evidence of “distillation”, which it suspects to be from DeepSeek. Given that many people have been using ChatGPT to distill their fine-tunes for a few years now, how can they be sure it was specifically DeepSeek? There's, say, glaive.ai whose entire business model is to sell you synthetic datasets, pr…
I didn't see this in o1 or any other LLM. Can distillation give deepseek such capability?