Live data from Hacker News

OpenAI says it has evidence DeepSeek used its model to train competitor

ft.com

701–710 of 1001 posts

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#701
post #61

> Furious [...] shocked I'm not seeing it. I get it, the narrative that OpenAI is getting a taste of their own medicine is funny but this is not serious reporting.

The link has been changed. My comment was about a different article that speculated on what OpenAI was "feeling" using hyperbole.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#702

Earlier quoted context omitted.

Why would it cast any doubt? If you can use o1 output to build a better R1. Then use R1 output to build a better X1... then a better X2.. XN, that just shows a method to create better systems for a fraction of the cost from where we stand. If it was that obvious OpenAI should have themselves done. But the disruptors did it. It hindsight it might sound obvious, but that is true for all innovations. It is all good stuf…

They're standing on the shoulders of giants, not only in terms of re-using expensive computing power almost for free by using the outputs of expensive models. It's a bit of a tradition in that country, also in manufacturing.

I thought OpenAI GPT took Wikipedia and the content of every book as inputs to train their models?

Everyone is standing on the shoulders of giants.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#703
post #676

The cat is out of the bag. This is the landscape now, r1 was made in a post-o1 world. Now other models can distill r1 and so on. I don’t buy the argument that distilling from o1 undermines deep seek’s claims around expense at all. Just as open AI used the tools ‘available to them’ to train their models (eg everyone else’ data), r1 is using today’s tools. Does open AI really have a moral or ethical high ground here?

Plus, it suggests OpenAI never had much of a moat.

Even if they win the legal case, it means weights can be inferred and improved upon simply by using the output that is also your core value add (e.g. the very output you need to sell to the world).

Their moat is about as strong as KFC's eleven herbs and spices. Maybe less...

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#704
This whole argument by OpenAI suggests they never had much of a moat.

Even if they win the legal case, it means weights can be inferred and improved upon simply by using the output that is also your core value add (e.g. the very output you need to sell to the world).

Their moat is about as strong as KFC's eleven herbs and spices. Maybe less...

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#705
post #320

I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…

This is going to have a catastrophic effect on closed source AI startup valuations. Because this means that anyone can copy any LLM. The person who trains the model, spends the most amount of money. Everyone else can create a replica at lower cost

Why is that bad? If a powerful entity can scrape every piece of media humanity has to offer and ignore copyright then why should society let then profit unrestricted from it? It's only fair that such models have no legal protection around their usage and can be used and analyzed by anyone as they see fit. The only reason this hasn't been codified into laws is because those same powerful entities have been busy trying to do regulatory capture.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#706

Earlier quoted context omitted.

> This is obviously extremely silly, because that's exactly how OpenAI got all of its training data IANAL, but It is worth noting here that DeepSeek has explicitly consented to a license that doesn't allow them to do this . That is a condition of using the Chat GPT and the OpenAI API. Even if the courts affirm that there's a fair use defence for AI training, DeepSeek may still be in the wrong here, not because of cop…

> DeepSeek has explicitly consented to a license that doesn't allow them to do this. By existing in USA, OpenAI consented to comply with copyright law, and how did that go?

[deleted]

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#707

Earlier quoted context omitted.

On another subject, if it belongs to OpenAI because it uses OpenAI, then doesn't that mean that everything produced using OpenAI belongs to OpenAI? Isn't that a reason not to use OpenAI? It's very similar to saying that you used Google and searched; now this product belongs to Google. They couldn't figure out how to respond; they went crazy.

to be clear, their terms of service are pretty clear that the USER owns the outputs.

The official stance in the US is currently that there is no copyright on AI output.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#708
post #320

I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…

> "DeepSeek trained on our outputs, and so their claims of replicating o1-level performance from scratch are not really true"

Someone has to correct me if I'm wrong, but I believe in ML research you always have a dataset and a model. They are distinct entities. It is plausible that output from OpenAI's model improved the quality of DeepSeek's dataset. Just like everyone publishing their code on GitHub improved the quality of OpenAI's dataset. What has been the thinking so far is that the dataset is not "part of" or "in" the model any more than the GPUs used to train the model are. It seems strange that that thinking should now change just because Chinese researchers did it better.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#710

Earlier quoted context omitted.

And seem to be more actively fulfilling the mission that 'Open'AI pretends to strive for.

Exactly, they actually opened up the model and research, which the "Open" company didn't, and merely adjusted some of their pricing tiers to try to combat commercially (but not without mumbling something like "yeah, we totally had these ideas too"). Now every single Meta, OpenAI etc engineer is trying to copy DeepSeek's innovations, and their first act is to... complain about copyright infringement, of all things?! W…

Altman was in a bit of a tricky position in that he figured OpenAI would need a lot of money for compute to be able to compete but it was hard to get that while remaining open. DeepSeek benefit from being funded from their own hedge fund. I wonder if part of their strategy is crack AI and then have it trade the markets?
Post reply on HN