I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…
This is going to have a catastrophic effect on closed source AI startup valuations. Because this means that anyone can copy any LLM. The person who trains the model, spends the most amount of money. Everyone else can create a replica at lower cost
OpenAI says it has evidence DeepSeek used its model to train competitor
671–680 of 1001 posts
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#672and this is good because any alternatives I can think of are older-school fighting
modern war is seeped in symbolism, but the contest is still there
e.g. whose dong is bigger? Xi Jingping's or Dnld Trump's
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#673Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#674I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#675I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…
Why would it cast any doubt? If you can use o1 output to build a better R1. Then use R1 output to build a better X1... then a better X2.. XN, that just shows a method to create better systems for a fraction of the cost from where we stand. If it was that obvious OpenAI should have themselves done. But the disruptors did it. It hindsight it might sound obvious, but that is true for all innovations. It is all good stuf…
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#676I don’t buy the argument that distilling from o1 undermines deep seek’s claims around expense at all. Just as open AI used the tools ‘available to them’ to train their models (eg everyone else’ data), r1 is using today’s tools.
Does open AI really have a moral or ethical high ground here?
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#677I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…
IANAL, but It is worth noting here that DeepSeek has explicitly consented to a license that doesn't allow them to do this. That is a condition of using the Chat GPT and the OpenAI API.
Even if the courts affirm that there's a fair use defence for AI training, DeepSeek may still be in the wrong here, not because of copyright infringement, but because of a breach of contract.
I don't think OpenAI would have much of a problem if you train your model on data scraped from the internet, some of which incidentally ends up being generated by Chat GPT.
Compare this to training AI models on Kindle Books randomly scraped off the internet, versus making a Kindle account, agreeing to the Kindle ToS, buying some books, breaking Amazon's DRM and then training your AI on that. What DeepSeek did is more analogous to the latter than the former.
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#678I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…
> This is obviously extremely silly, because that's exactly how OpenAI got all of its training data IANAL, but It is worth noting here that DeepSeek has explicitly consented to a license that doesn't allow them to do this . That is a condition of using the Chat GPT and the OpenAI API. Even if the courts affirm that there's a fair use defence for AI training, DeepSeek may still be in the wrong here, not because of cop…
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#679Earlier quoted context omitted.
> This is obviously extremely silly, because that's exactly how OpenAI got all of its training data IANAL, but It is worth noting here that DeepSeek has explicitly consented to a license that doesn't allow them to do this . That is a condition of using the Chat GPT and the OpenAI API. Even if the courts affirm that there's a fair use defence for AI training, DeepSeek may still be in the wrong here, not because of cop…
Did OpenAI abide by my service’s terms of service when it ingested my data?