Live data from Hacker News

OpenAI says it has evidence DeepSeek used its model to train competitor

ft.com

841–850 of 1001 posts

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#841
post #799
post #711

OpenAI is taking the position similar to that if you sell a cook book, people are not allowed to teach the recipes to their kids, or make better versions of them. That is absurd. Copyright law is designed to strike a balance between two issues. One the one hand, the creator’s personality that’s baked into the specific form of expression. And on the other hand, society’s interest in ideas being circulated, improved an…

The stuff about copyright seems irrelevant. OpenAI's future investments -- billions -- were just threatened to be undercut by several orders of magnitude by a competitor. It's in their best interests to cast doubt on that competitor's achievements. If they can do so by implying that OpenAI are in fact the source of most of the DeepSeek's performance then all the better. It doesn't matter whether there's a compelling…

>It just needs to be plausible enough that OpenAI can make a reasonable case for continuing investment at the levels it's historically attained

>there's now doubt that a DeepSeek-level model could be trained without making use of OpenAI's substantial levels of investment.

But, this still seems to be a problem for OpenAI. Who wants to invest "substantially" in a company whose output can be used by competitors to build an equal or better offering for orders of magnitude less?

Seems they'd need to make that copyright stick. But, that's a very tall and ironic order, given how OpenAI obtained its data in the first place.

There's a scenario where this development is catastrophic for OpenAI's business model.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#842
post #819

Earlier quoted context omitted.

TikTok is banned in the US?

Yes, it was removed from the app stores, and briefly, from the web.

Except access to the app didn't have to stop. TikTok chose to manipulate users and Trump by going beyond the law and kissing Trump's rear. It was only US companies that couldn't host the app (eg Google and Apple). Users in the US could have still accessed the app, and even side-loaded it on Android, but TikTok purposely blocked them and pretended it was the ban. They were able to do it because they know the exact location of every TikTok user whether you use a VPN or not.

Source:

> If not sold within a year, the law would make it illegal for web-hosting services to support TikTok, and it would force Google and Apple to remove TikTok from app stores — rendering the app unusable with time.

https://www.npr.org/2024/04/24/1246663779/biden-ban-tiktok-u...

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#843
post #799
post #711

OpenAI is taking the position similar to that if you sell a cook book, people are not allowed to teach the recipes to their kids, or make better versions of them. That is absurd. Copyright law is designed to strike a balance between two issues. One the one hand, the creator’s personality that’s baked into the specific form of expression. And on the other hand, society’s interest in ideas being circulated, improved an…

The stuff about copyright seems irrelevant. OpenAI's future investments -- billions -- were just threatened to be undercut by several orders of magnitude by a competitor. It's in their best interests to cast doubt on that competitor's achievements. If they can do so by implying that OpenAI are in fact the source of most of the DeepSeek's performance then all the better. It doesn't matter whether there's a compelling…

Isn’t this precisely how so many opensource LLMs caught up with OpenAI so quickly, because they could just train on actual ChatGPT output?

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#844

Earlier quoted context omitted.

They're standing on the shoulders of giants, not only in terms of re-using expensive computing power almost for free by using the outputs of expensive models. It's a bit of a tradition in that country, also in manufacturing.

I thought OpenAI GPT took Wikipedia and the content of every book as inputs to train their models? Everyone is standing on the shoulders of giants.

What I meant to say was that OpenAI did put a lot of money into extracting value out of the pile of (partially copyrighted) data, and that DeepSeek was freeloading on that investment without disclosing it, making them look more efficient than they truly are.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#845
post #799
post #711

OpenAI is taking the position similar to that if you sell a cook book, people are not allowed to teach the recipes to their kids, or make better versions of them. That is absurd. Copyright law is designed to strike a balance between two issues. One the one hand, the creator’s personality that’s baked into the specific form of expression. And on the other hand, society’s interest in ideas being circulated, improved an…

The stuff about copyright seems irrelevant. OpenAI's future investments -- billions -- were just threatened to be undercut by several orders of magnitude by a competitor. It's in their best interests to cast doubt on that competitor's achievements. If they can do so by implying that OpenAI are in fact the source of most of the DeepSeek's performance then all the better. It doesn't matter whether there's a compelling…

OpenAI is going out of their way to demonstrate that they will willingly spend the money of their investors to the tune of 100s of billions of dollars, only to then enable 100s of derivative competitors that can be launched at a fraction of the cost.

Basically, in a round about way, OpenAi is going back to their roots and more - they're something between a charity and Robin Hood, stealing the money of rich investors and giving it to poor and aspirational AI competitors.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#846
post #320

I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…

That's still problematic because any model that OpenAI trains can now be "stolen" and essentially rendered "open".

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#847

Earlier quoted context omitted.

You can't copyright AI generated works. OpenAI are barking up the wrong tree.

They're not making a legal claim, they're trying to establish provenance over Deepseek in the public eye.

The public gives no shit about any of these companies. There are no moats or loyalties in this space. Just individuals and corporations looking for the cheapest tool possible that gets the job close to done.

OpenAi spent investor money to enable random Chinese Ai startups to offer a better version of their own product at a fraction of the cost. In some ways, this was inevitable to be the conclusion, but I do find the way we arrive at this conclusion to be particularly enjoyable to watch playout.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#848
post #454

Earlier quoted context omitted.

I'm willing to bet ''ban DeepSeek'' voices will start soon. Why compete, when you can just ban?

They've started already, I've seen posts on LinkedIn implying or outright stating that DeepSeek is a national security risk (IMHO, LinkedIn being the social media outlet most corporate-sycophantic). I went ahead and just picked this one at random from my feed. https://www.linkedin.com/posts/kevinkeller_deepseek-privacy-...

At least this guy can differentiate between running your own model and using the web/mobile app where DeepSeek process your data. I've watched a TV show yesterday (I think it was France24) where the "experts" can't really tell the difference or are not aware of it. Shut down the TV and went to sleep.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#850
post #320

I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…

> "DeepSeek trained on our outputs"

I'm wondering how Deepseek could have made 100s of millions of training queries to OpenAI and not one person at OpenAI caught on.

Post reply on HN