I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…
There is a third possibility I haven't seen discussed yet: That DeepSeek, illegally, got their hands on an OpenAI model via a breach of OpenAI's systems. Its easy to laugh at OpenAI and say "you reap what you sow", I'm 100% in that camp, but given the lengths other Chinese entities have gone to when it comes to replicating Western technology; we should not discount this. That being said, breaching OAI's systems, re-t…
OpenAI says it has evidence DeepSeek used its model to train competitor
581–590 of 1001 posts
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#582Earlier quoted context omitted.
I think it would cast doubt on the narrative "you could have trained o1 with much less compute, and r1 is proof of that", if it turned out that in order to train r1 in the first place, you had to have access to bunch of outputs from o1. In other words, you had to do the really expensive o1 training in the first place. (with the caveat that all we have right now are accusations that DeepSeek made use of OpenAI data -…
o1 wouldn't exist without the combined compute of every mind that led to the training data they used in the first place. How many h100 equivalents are the rolling continuum of all of human history?
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#583I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…
Even if they didn’t directly, intentionally use o1 output (and they didn’t claim they didn’t, so far as I know), AI slop is everywhere. We passed peak original content years ago. Everything is tainted and everything should be understand in that context.
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#584All the top level comments are basking in the irony of it, which is fair enough. But I think this changes the Deepseek narrative a bit. If they just benefited from repurposing OpenAI data, that's different than having achieved an engineering breakthrough, which may suggest OpenAI's results were hard earned after all.
These aren't mutually exclusive. It's been known for a while that competitors used OpenAI to improve their models, that's why they changed the TOS to forbid it. That doesn't mean the deep seek technical achievements are less valid.
Well, that's literally exactly what it would mean. If DeepSeek relied on OpenAI’s API, their main achievement is in efficiency and cost reduction as opposed to fundamental AI breakthroughs.
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#585I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…
There is a third possibility I haven't seen discussed yet: That DeepSeek, illegally, got their hands on an OpenAI model via a breach of OpenAI's systems. Its easy to laugh at OpenAI and say "you reap what you sow", I'm 100% in that camp, but given the lengths other Chinese entities have gone to when it comes to replicating Western technology; we should not discount this. That being said, breaching OAI's systems, re-t…
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#586I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…
There is a third possibility I haven't seen discussed yet: That DeepSeek, illegally, got their hands on an OpenAI model via a breach of OpenAI's systems. Its easy to laugh at OpenAI and say "you reap what you sow", I'm 100% in that camp, but given the lengths other Chinese entities have gone to when it comes to replicating Western technology; we should not discount this. That being said, breaching OAI's systems, re-t…
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#587Earlier quoted context omitted.
There is a third possibility I haven't seen discussed yet: That DeepSeek, illegally, got their hands on an OpenAI model via a breach of OpenAI's systems. Its easy to laugh at OpenAI and say "you reap what you sow", I'm 100% in that camp, but given the lengths other Chinese entities have gone to when it comes to replicating Western technology; we should not discount this. That being said, breaching OAI's systems, re-t…
The reason you’re not seeing that being discussed is it’s totally unsupported by any evidence that’s in the public domain. Unless you have some actual evidence of such a breach, you may as well introduce the possibility that DeepSeek was reverse engineered from data found at an alien crash site.
The Chinese Communist party very much sees itself in a global rivalry over "new productive forces". That's official policy. And US leadership basically agrees.
The US is playing dirty by essentially embargoing China over big AI - why wouldn't it occur to them to retaliate by playing dirtier?
I mean we probably won't know for sure, but it's much less far fetched than a lot of other speculation in this area.
E.g., R1's cold start training could probably have benefited quite a bit from having access to OpenAI's chain of thought data for training. The paper is a bit light on detail on how it was made.
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#588Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#589Earlier quoted context omitted.
I think the point is - OpenAI scraped public data - d1 - Trained their model to produce output - d2 - DeepSeek used d2 to reinforce their model OpenAI is mad about d2 (not d1). I'm not sure using public data is "stealing". In summary, these are two different things & need to be separate.
You say "public", but what I think you mean is "publicly available". Even publicly available data has copyrights, and unless that copyright is "public domain", you need to follow some rules. Even licenses like Creative Commons, which would be the most permissive, come with caveats which OpenAI doesn't follow [0]. It is unclear if someone breaking someone else's copyright to use A can claim copyright on a work B, deri…
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#590All the top level comments are basking in the irony of it, which is fair enough. But I think this changes the Deepseek narrative a bit. If they just benefited from repurposing OpenAI data, that's different than having achieved an engineering breakthrough, which may suggest OpenAI's results were hard earned after all.
I really don't see a correlation here to be honest. Eventually all future AIs will be produced with synthetic input, the amount of (quality) data we humans can produce is quite limited. The fact that the input of one AI has been used in the training of another one seems irrelevant.
The deeper question is whether Deepseek has achieved real autonomy or if it’s just a derivative work. If the latter, then OpenAI still holds the keys to future advances. If Deepseek truly found a way to be independent while achieving similar performance, then OpenAI has a problem.
The details of how they trained matter more than the inevitability of synthetic data down the line.