Live data from Hacker News

OpenAI says it has evidence DeepSeek used its model to train competitor

ft.com

551–560 of 1001 posts

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#551
I was wondering if this might be the case, similar to how Bing’s initial training included Google’s search results [1]. I’d be curious to see more details of OpenAI’s evidence.

It is, of course, quite ironic for OpenAI to indiscriminately scrape the entire web and then complain about being scraped themselves.

[1]: https://searchengineland.com/google-bing-is-cheating-copying...

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#552

I do think that distilling a model from another is much less impressive than distilling one from raw text. However, it is hard to say if it is really illegal or even immoral, perhaps just one step further in the evolution of the space.

Isn't it more impressive given that training on model output usually leads to worse model?

If they actually figured out how to use output of existing models to build model that outperforms them then it's something that brings us closer to singularity than every other development so far.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#553
post #125

> “It is (relatively) easy to copy something that you know works,” Altman tweeted. “It is extremely hard to do something new, risky, and difficult when you don’t know if it will work.” The humor/hypocrisy of the situation aside, it does seem to be true that OpenAI is consistently the one coming up with new ideas first (GPT 4, o1, 4o-style multimodality, voice chat, DALL-E, …) and then other companies reproduce their…

> other companies reproduce their work, and get more credit because they actually publish the research. I don't understand, you mean OpenAI isn't releasing open models and openly publishing their research?

Are you being sarcastic (honestly, it's hard to tell after reading as many uninformed takes in the past week as I have).

No, they aren't (other than whisper).

Their "papers" are closer to marketing materials. Very intentionally leaving out tons of technical information.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#554
I mean, if openAI claims they can train on the world’s novels and blogs with “no harm done” (i.e: no copyright infringement and no royalties due), then it directly follows that we can train both our robots and our selves on the output of openAI’s models in kind.

Right?

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#555

I think OpenAI is in a really weak position here. There are essentially two positions you can be in: You can be the agile new startup that can break the rules and move fast. That's what OpenAI used to be. Or you can be the big incumbent who is going to use your enormous resources to crush your opposition. That's Google & Microsoft here. For Microsoft to say "We're going to tie you up in lawsuits about the way you tra…

But microsoft is one of their backers?

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#556

All the top level comments are basking in the irony of it, which is fair enough. But I think this changes the Deepseek narrative a bit. If they just benefited from repurposing OpenAI data, that's different than having achieved an engineering breakthrough, which may suggest OpenAI's results were hard earned after all.

> OpenAI's results were hard earned after all DDOSing web sites and grabbing content without anyone's consent is not hard earned at all. They did spent billions on their thing, but nothing was earned as they could never do that legally.

I understand the temptation to go there, but I think it misses the point. I have no qualms at all with the idea that the sum total of intelligence distributed across the internet was siphoned away from creators and piped through an engine that now cynically seeks to replace them. Believe me, I will grab my pitchfork and march side by side with you.

But let's keep the eye on the ball for a second. None of that changes the fact that what was built was a capability to reflect that knowledge in dynamic and deep ways in conversation, as well as image and audio recognition.

And did Deepseek also build that? From scratch? Because they might not have.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#557
post #320

I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…

Why would it cast any doubt? If you can use o1 output to build a better R1. Then use R1 output to build a better X1... then a better X2.. XN, that just shows a method to create better systems for a fraction of the cost from where we stand. If it was that obvious OpenAI should have themselves done. But the disruptors did it. It hindsight it might sound obvious, but that is true for all innovations. It is all good stuf…

"Then use R1 output to build a better X1" is the part I'm not sure about. Is X1 going to actually be better than R1?

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#558
post #320

I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…

On another subject, if it belongs to OpenAI because it uses OpenAI, then doesn't that mean that everything produced using OpenAI belongs to OpenAI? Isn't that a reason not to use OpenAI? It's very similar to saying that you used Google and searched; now this product belongs to Google. They couldn't figure out how to respond; they went crazy.

The US ruled that AI produced things are by themself not copyrightable.

So no, it doesn't belong to OpenAI.

You might be able to sue for penalties for breach of contract of the TOS, but that doesn't give them the right to the model. And even if it doesn't give them any right to invalidate unbound copyright grants they have given to 3rd parties (here literally everyone). Nor does it prevent anyone from training their own new models based on it or prevent anyone from using it. Oh, and the one breaching the TOS might not even have been the company behind DeepSeek but some in-between 3rd party.

Naturally this is under a few assumptions:

- the US consistently applies it's own law, but they have a long history of not doing so

- the US doesn't abuse their power to force their economical opinions (ban DeepSeek) on other countries

- it actually was trained on OpenAI, but uh, OpenAI has IMHO shown over the years very clearly that they can't be trusted and they are fully in-transparent. How do we trust their claim? How do we trust them to not retrospectively have tweaked their model to make it look as if DeepSeek copied it?

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#559
post #500

Earlier quoted context omitted.

Why would it cast any doubt? If you can use o1 output to build a better R1. Then use R1 output to build a better X1... then a better X2.. XN, that just shows a method to create better systems for a fraction of the cost from where we stand. If it was that obvious OpenAI should have themselves done. But the disruptors did it. It hindsight it might sound obvious, but that is true for all innovations. It is all good stuf…

Why not just copy and paste the model and change the name? That's an even more efficient form of distillation.

Even assuming the model was somehow publicly available in a form that could be directly copied, that would be a more blatant form of copyright infringement. Distillation launders copyrighted material in a way that OpenAI specifically has argued falls under fair use.
Post reply on HN