Live data from Hacker News

OpenAI says it has evidence DeepSeek used its model to train competitor

ft.com

601–610 of 1001 posts

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#601
post #320

I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…

You're right that the first claim is silly, but the second claim is pretty silly too — they're not claiming industrial espionage, they're claiming a breach in ToS. The outputs of the o1 thinking process aren't user-visible, and never leave OpenAI's datacenters. Unless DeepSeek actually had a mole that stole their o1 outputs, there's nothing useful DeepSeek could've distilled to get to R1's thought processes.

And if DeepSeek had a mole, why would they bother running a massive job internally to steal the data generated? It would be way easier for the mole to just leak the RL training process, and DeepSeek could quietly copy it rather than bothering with exfiltrating massive datasets to distill. The training process is most likely like, on the order of a hundred lines of Python or so, and you don't even need the file: you just need someone to describe it to you. Much simpler than snatching hundreds of gigabytes of training data off of internal servers...

Plus, the RL process described in DeepSeek's paper has already been replicated by a PhD student at Berkeley: https://x.com/karpathy/status/1884678601704169965 So, it seems pretty unlikely they simply distilled R1 and lied about it, or else how does their RL training algo actually... work?

This is mainly cope from OpenAI that their supposedly super duper advanced models got caught by China within a few months of release, for way cheaper than it cost OpenAI to train.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#603
post #465
post #320

I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…

There is a third possibility I haven't seen discussed yet: That DeepSeek, illegally, got their hands on an OpenAI model via a breach of OpenAI's systems. Its easy to laugh at OpenAI and say "you reap what you sow", I'm 100% in that camp, but given the lengths other Chinese entities have gone to when it comes to replicating Western technology; we should not discount this. That being said, breaching OAI's systems, re-t…

Can you explain at a technical level how you view this as necessary for the observed result?

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#604

Earlier quoted context omitted.

I really don't see a correlation here to be honest. Eventually all future AIs will be produced with synthetic input, the amount of (quality) data we humans can produce is quite limited. The fact that the input of one AI has been used in the training of another one seems irrelevant.

The issue isn’t just that AI trained on AI is inevitable it's whose AI is being used as the base layer. Right now, OpenAI’s models are at the top of that hierarchy. If Deepseek depended on them, it means OpenAI is still the upstream bottleneck, not easily replaced. The deeper question is whether Deepseek has achieved real autonomy or if it’s just a derivative work. If the latter, then OpenAI still holds the keys to f…

> then OpenAI still holds the keys to future advances

Point is, those future advances are worthless. Eventually anybody will be able to feed each other's data for the training.

There's no moat here. LLMs are commodities.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#605
post #600
post #320

I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…

> This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. OpenAI has also invested heavily in human annotation and RLHF. If all DeepSeek wanted was a proxy for scraped training data, they'd probably just scrape it themselves. Using existing RLHF'd models as replacement for expensive humans in the training…

"We spent a lot of labor processing everything we stole" is...not how that works.

That's like the mafia complaining that they worked so hard to steal those barrels of beer that someone made off with in the middle of the night and really that's not fair and won't someone do something about it?

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#607
post #87

I'm not being sarcastic, but we may soon have to torrent DeepSeek's model. OpenAI has a lot of clout in the US and could get DeepSeek banned in western countries for copyright.

that would be suicide - that company only exists because they stole content for every single person, website and media company on the planet.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#608
post #504

Earlier quoted context omitted.

I think it would cast doubt on the narrative "you could have trained o1 with much less compute, and r1 is proof of that", if it turned out that in order to train r1 in the first place, you had to have access to bunch of outputs from o1. In other words, you had to do the really expensive o1 training in the first place. (with the caveat that all we have right now are accusations that DeepSeek made use of OpenAI data -…

From the R1 paper In this study, we demonstrate that reasoning capabilities can be significantly improved through large-scale reinforcement learning (RL), even without using supervised fine-tuning (SFT) as a cold start. Furthermore, performance can be further enhanced with the inclusion of a small amount of cold-start data Is this cold start data what OpenAI is claiming their output ? If so what's the big deal ?

DeepSeek claims that the cold-start data is from DeepSeekV3, which is the model that has the $5.5M pricetag. If that data were actually the output of o1 (a model that had a much higher training cost, and its own RL post-training), that would significantly change the narrative of R1's development, and what's possible to build from scratch on a comparable training budget.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#609
post #4

How would they prove they used it’s model. I would be curious to know their methodology. Also what legal actions OpenAI can take? can DeepSeek be banned in US?

They might show DeepSeek's model calling itself ChatGPT, which users have already alleged. Same as how Cisco proved Huawei was stealing router code. Except in this case, nothing was stolen, unless they want to call ChatGPT's own training on source data theft too.

ChatGPT outputs are all over the internet. It is harder to prove that deepseek used specifically o1 for training, instead of a lot of chatgpt output ending up in the training set from other sources.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#610

Earlier quoted context omitted.

> other companies reproduce their work, and get more credit because they actually publish the research. I don't understand, you mean OpenAI isn't releasing open models and openly publishing their research?

Are you being sarcastic (honestly, it's hard to tell after reading as many uninformed takes in the past week as I have). No, they aren't (other than whisper). Their "papers" are closer to marketing materials. Very intentionally leaving out tons of technical information.

They are being sarcastic.
Post reply on HN