Live data from Hacker News

OpenAI says it has evidence DeepSeek used its model to train competitor

ft.com

491–500 of 1001 posts

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#491
post #320

I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…

Why would it cast any doubt? If you can use o1 output to build a better R1. Then use R1 output to build a better X1... then a better X2.. XN, that just shows a method to create better systems for a fraction of the cost from where we stand. If it was that obvious OpenAI should have themselves done. But the disruptors did it. It hindsight it might sound obvious, but that is true for all innovations. It is all good stuf…

OpenAI couldn't do it, when the high cost of training and access to GPUs is their competitive advance against startups, they can't admit that it does not exist.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#492
post #431

[flagged]

And we'd be shooting ourselves in the foot to do so. If America is forced to use only the clunky corporate-owned American AI at a fee, we'll very quickly fall behind competitors worldwide who use DeepSeek models to produce better results for much, much cheaper.

Not to mention it'd defeat the whole purpose of a "free market" economy. (Not that that means much of anything anymore)

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#493
post #374

Everyone is responding to the intellectual property issue, but isn't that the less interesting point? If Deepseek trained off OpenAI, then it wasn't trained from scratch for "pennies on the dollar" and isn't the Sputnik-like technical breakthrough that we've been hearing so much about. That's the news here. Or rather, the potential news, since we don't know if it's true yet.

Even if all that about training is true, the bigger cost is inference and Deepseek is 100x cheaper. That destroys OpenAI/Anthropic's value proposition of having a unique secret sauce so users are quickly fleeing to cheaper alternatives.

Google Deepmind's recent Gemini 2.0 Flash Thinking is also priced at the new Deepseek level. It's pretty good (unlike previous Gemini models).

[0] https://x.com/deedydas/status/1883355957838897409

[1] https://x.com/raveeshbhalla/status/1883380722645512275

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#494
post #465

Earlier quoted context omitted.

There is a third possibility I haven't seen discussed yet: That DeepSeek, illegally, got their hands on an OpenAI model via a breach of OpenAI's systems. Its easy to laugh at OpenAI and say "you reap what you sow", I'm 100% in that camp, but given the lengths other Chinese entities have gone to when it comes to replicating Western technology; we should not discount this. That being said, breaching OAI's systems, re-t…

The reason you’re not seeing that being discussed is it’s totally unsupported by any evidence that’s in the public domain. Unless you have some actual evidence of such a breach, you may as well introduce the possibility that DeepSeek was reverse engineered from data found at an alien crash site.

[flagged]

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#495
post #465

Earlier quoted context omitted.

There is a third possibility I haven't seen discussed yet: That DeepSeek, illegally, got their hands on an OpenAI model via a breach of OpenAI's systems. Its easy to laugh at OpenAI and say "you reap what you sow", I'm 100% in that camp, but given the lengths other Chinese entities have gone to when it comes to replicating Western technology; we should not discount this. That being said, breaching OAI's systems, re-t…

[flagged]

> based purely on racial prejudices

I don't think that's what the parent was getting at. The US and China are in an ongoing "cyber war". Both sides of that conflict actively use their computers to send messages/signals to other computers, hoping that the exploits contained in those messages/signals can be used to exfiltrate data from and/or gain control of the computer receiving the message. It would really be weird to flatly discount the possibility that some OpenAI data was leaked, however closely guarded it may be.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#496
post #465
post #320

I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…

There is a third possibility I haven't seen discussed yet: That DeepSeek, illegally, got their hands on an OpenAI model via a breach of OpenAI's systems. Its easy to laugh at OpenAI and say "you reap what you sow", I'm 100% in that camp, but given the lengths other Chinese entities have gone to when it comes to replicating Western technology; we should not discount this. That being said, breaching OAI's systems, re-t…

[flagged]

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#497
post #465

Earlier quoted context omitted.

There is a third possibility I haven't seen discussed yet: That DeepSeek, illegally, got their hands on an OpenAI model via a breach of OpenAI's systems. Its easy to laugh at OpenAI and say "you reap what you sow", I'm 100% in that camp, but given the lengths other Chinese entities have gone to when it comes to replicating Western technology; we should not discount this. That being said, breaching OAI's systems, re-t…

[flagged]

[flagged]

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#498
post #442
post #125

> “It is (relatively) easy to copy something that you know works,” Altman tweeted. “It is extremely hard to do something new, risky, and difficult when you don’t know if it will work.” The humor/hypocrisy of the situation aside, it does seem to be true that OpenAI is consistently the one coming up with new ideas first (GPT 4, o1, 4o-style multimodality, voice chat, DALL-E, …) and then other companies reproduce their…

Boy who stole test papers complains about child copying his answers.

No you don’t understand, AI is “dangerous” and only him and his uber rich billionaire mates should get to control it!

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#499
post #95

> “It’s also extremely hard to rally a big talented research team to charge a new hill in the fog together,” he added. “This is the key to driving progress forward.” Well I think DeepSeek releasing it open source and on an MIT license will rally the big talent. The open sourcing of a new technology has always driven progress in the past. The last paragraph too is where OpenAi seems to be focusing their efforts.. > we…

> So they'll go for getting DeepSeek banned like TikTok was now that a precedent has been set ? Can't really ban what can be downloaded for free and hosted by anyone. There are many providers hosting the ~700B parameter version that aren't CCP aligned.

I'm old enough to remember when the US government did something very similar. For years (decades?), we banned any implementation of public-key cryptography under the guise of the technology being akin to munitions.

People made shirts with printouts of the code to RSA under the heading "this shirt is a munition." Apparently such shirts are still for sale, even though they are not classified as munitions anymore.

[1] - https://en.wikipedia.org/wiki/Export_of_cryptography_from_th...

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#500
post #320

I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…

Why would it cast any doubt? If you can use o1 output to build a better R1. Then use R1 output to build a better X1... then a better X2.. XN, that just shows a method to create better systems for a fraction of the cost from where we stand. If it was that obvious OpenAI should have themselves done. But the disruptors did it. It hindsight it might sound obvious, but that is true for all innovations. It is all good stuf…

Why not just copy and paste the model and change the name? That's an even more efficient form of distillation.
Post reply on HN