Live data from Hacker News

OpenAI says it has evidence DeepSeek used its model to train competitor

ft.com

671–680 of 1001 posts

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#671
post #320

I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…

This is going to have a catastrophic effect on closed source AI startup valuations. Because this means that anyone can copy any LLM. The person who trains the model, spends the most amount of money. Everyone else can create a replica at lower cost

Maybe anyone can copy any LLM with sufficient querying. There are still ways to guard one.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#672
I see this as China fighting U.S. of A (or the American Dollar versus Chinese Renmibi if you will)

and this is good because any alternatives I can think of are older-school fighting

modern war is seeped in symbolism, but the contest is still there

e.g. whose dong is bigger? Xi Jingping's or Dnld Trump's

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#674
post #320

I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…

I think the more interesting claim (that Deepseek should make for lols) is that it wasn't them who trained R1. No, it was O1's idea. It chose to take the young R1 as its padawan.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#675
post #320

I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…

Why would it cast any doubt? If you can use o1 output to build a better R1. Then use R1 output to build a better X1... then a better X2.. XN, that just shows a method to create better systems for a fraction of the cost from where we stand. If it was that obvious OpenAI should have themselves done. But the disruptors did it. It hindsight it might sound obvious, but that is true for all innovations. It is all good stuf…

They did do that themselves it's called o3.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#676
The cat is out of the bag. This is the landscape now, r1 was made in a post-o1 world. Now other models can distill r1 and so on.

I don’t buy the argument that distilling from o1 undermines deep seek’s claims around expense at all. Just as open AI used the tools ‘available to them’ to train their models (eg everyone else’ data), r1 is using today’s tools.

Does open AI really have a moral or ethical high ground here?

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#677
post #320

I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…

> This is obviously extremely silly, because that's exactly how OpenAI got all of its training data

IANAL, but It is worth noting here that DeepSeek has explicitly consented to a license that doesn't allow them to do this. That is a condition of using the Chat GPT and the OpenAI API.

Even if the courts affirm that there's a fair use defence for AI training, DeepSeek may still be in the wrong here, not because of copyright infringement, but because of a breach of contract.

I don't think OpenAI would have much of a problem if you train your model on data scraped from the internet, some of which incidentally ends up being generated by Chat GPT.

Compare this to training AI models on Kindle Books randomly scraped off the internet, versus making a Kindle account, agreeing to the Kindle ToS, buying some books, breaking Amazon's DRM and then training your AI on that. What DeepSeek did is more analogous to the latter than the former.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#678
post #320

I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…

> This is obviously extremely silly, because that's exactly how OpenAI got all of its training data IANAL, but It is worth noting here that DeepSeek has explicitly consented to a license that doesn't allow them to do this . That is a condition of using the Chat GPT and the OpenAI API. Even if the courts affirm that there's a fair use defence for AI training, DeepSeek may still be in the wrong here, not because of cop…

Did OpenAI abide by my service’s terms of service when it ingested my data?

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#679
post #678

Earlier quoted context omitted.

> This is obviously extremely silly, because that's exactly how OpenAI got all of its training data IANAL, but It is worth noting here that DeepSeek has explicitly consented to a license that doesn't allow them to do this . That is a condition of using the Chat GPT and the OpenAI API. Even if the courts affirm that there's a fair use defence for AI training, DeepSeek may still be in the wrong here, not because of cop…

Did OpenAI abide by my service’s terms of service when it ingested my data?

Did OpenAI have to sign up for your service to gain access?

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#680

[flagged]

[flagged]

What double standard? I have not claimed anything about openAI being ethical paragon of virtue.

I anything, you can be accused of using whataboutism in order to justify DeepSeek illegal actions

Post reply on HN