Live data from Hacker News

OpenAI says it has evidence DeepSeek used its model to train competitor

ft.com

731–740 of 1001 posts

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#731
post #320

I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…

> DeepSeek trained on our outputs, and so their claims of replicating o1-level performance from scratch are not really true" This is at least plausibly a valid claim.

Some may view this as partially true, given that o-1 does not output its CoT process.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#732
post #711

OpenAI is taking the position similar to that if you sell a cook book, people are not allowed to teach the recipes to their kids, or make better versions of them. That is absurd. Copyright law is designed to strike a balance between two issues. One the one hand, the creator’s personality that’s baked into the specific form of expression. And on the other hand, society’s interest in ideas being circulated, improved an…

Just to play devil's advocate, OAI can argue that they spent great effort creating and procuring annotated data. Such datasets are indeed their secret, and now DS gets them for free by distilling OAI's output. Besides, OAI's EULA explicitly forbids users from using the output of their API for model training. I'm not saying that OAI is right, of course. Just to present OAI's point of view.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#734
Boo hoo?

Back in college, a kid in my dorm had a huge MP3 collection. And he shared it out over the network, and people were all like, "Man, Patrick has an amazing MP3 collection!" And he spent hours and hours ripping CDs from everyone so all the music was available on our network.

Then I remember another kid coming in, with a bigger hard drive, and he just copied all of Patrick's MP3 collection and added a few more to it. Then ran the whole thing through iTunes to clean up names and add album covers. It was so cool!

And I remember Patrick complained, "He stole my MP3 collection!"

Anyway this story sums up how I feel about Sam Altman here. He's not Metalica, he's Patrick.

https://www.npr.org/2023/12/27/1221821750/new-york-times-sue...

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#735

I think OpenAI is in a really weak position here. There are essentially two positions you can be in: You can be the agile new startup that can break the rules and move fast. That's what OpenAI used to be. Or you can be the big incumbent who is going to use your enormous resources to crush your opposition. That's Google & Microsoft here. For Microsoft to say "We're going to tie you up in lawsuits about the way you tra…

> they're still a minnow 3K+ employees, $3B+ revenue, ... sure, not BigTech but hardly a minnow. A company that big can chew gum and walk at the same time.

They're trying to bark up a tree that might happen to be backed by the People's Republic of China. That's not their league, and even Microsoft would think twice before getting into that kind of kerfuffle.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#736

Earlier quoted context omitted.

> This is obviously extremely silly, because that's exactly how OpenAI got all of its training data IANAL, but It is worth noting here that DeepSeek has explicitly consented to a license that doesn't allow them to do this . That is a condition of using the Chat GPT and the OpenAI API. Even if the courts affirm that there's a fair use defence for AI training, DeepSeek may still be in the wrong here, not because of cop…

training is either fair use, or it isn't OpenAI can't have it both ways

Right, but it was never about doing the right thing for humanity, it was about doing the right thing for their profits.

Like I’ve said time and time again, nobody in this space gives a fuck about anyone that isn’t directly contributing money to their bottom line at that particular instant. The fundamental idea is selfish, damages the fundamental machinery that makes the internet useful by penalizing people that actually make things, and will never, ever do anything for the greater good if it even stands a chance of reducing their standing in this ridiculously overhyped market. Giving people free access to what is for all intents and purposes a black box is not “open” anything, is no more free (as in speech) than Slack is, and all of this is obviously them selling a product at a huge loss to put competing media out of business and grab market share.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#738
post #711

OpenAI is taking the position similar to that if you sell a cook book, people are not allowed to teach the recipes to their kids, or make better versions of them. That is absurd. Copyright law is designed to strike a balance between two issues. One the one hand, the creator’s personality that’s baked into the specific form of expression. And on the other hand, society’s interest in ideas being circulated, improved an…

You can't copyright AI generated works. OpenAI are barking up the wrong tree.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#739
post #320

I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…

Even for the latter point (If true, I'd call this assertion highly questionable), so what?

That's honestly such a academic point, who really cares?

They've been outcompeted and the argument is 'well if we didn't let people access our models, they would of taken longer to get here' so what??

The only thing this gets them is an explanation as to why training o1 cost them more than 5 million or whatever, but that is in the past the datacentre has consumed the energy.. the money has gone up in fairly literal steam.

Post reply on HN