Live data from Hacker News

OpenAI says it has evidence DeepSeek used its model to train competitor

ft.com

761–770 of 1001 posts

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#761
post #320

I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…

It’s literally a race to the bottom by “theft of data”

Whatever that means. The legal system right now in shambles and flat footed.

Knowing our current government leadership, I think we’re going to see some brute force action backed up by the United States military.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#762
If the material which OpenAI is trained on is itself not subject to copyright protections, then other LLMs trained on OpenAI should also not be subject to any copyright restrictions.

You can't have both ways... If OpenAI wants to claim that the AI is not repeating content but 'synthesizing it' in the same was as a human student would do... Then I think the same logic should extend to DeepSeek.

Now if OpenAI wants to claim that its own output is in fact copyright-protected, then it seems like it should owe royalty payments to everyone whose content was sourced upstream to build its own training set. Also, synthetic content which is derived from real content should also be factored in.

TBH, this could make a strong case for taxing AI. Like some kind of fee for human knowledge and distributed as UBI. The training data played a key part in this AI innovation.

As an open source coder, I know that my copyrighted code is being used by AI to help other people produce derived code and, by adapting it in this way, it's making my own code less relevant to some extent... In effect, it could be said that my code has been mixed in with the code of other open source developers and weaponized against us.

It feels like it could go either way TBH but there needs to be consistency.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#763

Earlier quoted context omitted.

You can't copyright AI generated works. OpenAI are barking up the wrong tree.

They're not making a legal claim, they're trying to establish provenance over Deepseek in the public eye.

If this is their goal, R1 is on par with the $200 a month model. Most people don't give a shit.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#764
post #743

Earlier quoted context omitted.

Just to play devil's advocate, OAI can argue that they spent great effort creating and procuring annotated data. Such datasets are indeed their secret, and now DS gets them for free by distilling OAI's output. Besides, OAI's EULA explicitly forbids users from using the output of their API for model training. I'm not saying that OAI is right, of course. Just to present OAI's point of view.

So a bank robber that manages to steal from Fort Knox gets to keep the gold bars because it was a very complicated job?

If the Fort Knox gold was originally stolen from the Incas

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#765
post #95

> “It’s also extremely hard to rally a big talented research team to charge a new hill in the fog together,” he added. “This is the key to driving progress forward.” Well I think DeepSeek releasing it open source and on an MIT license will rally the big talent. The open sourcing of a new technology has always driven progress in the past. The last paragraph too is where OpenAi seems to be focusing their efforts.. > we…

The fact that they are still called "Open"AI adds such a delicious irony to this whole thing. I could not imagine a company I had less sympathy for in this situation.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#766

Earlier quoted context omitted.

[flagged]

BYD cars are everywhere in Latin America and Europe. Xiaomi phones also.

I live in Europe half the time. Literally never seen a BYD car in the flesh.

Chinese phones, yes. But I'd argue we're past peak China. Huawei phones briefly were the #1 selling in the world, have since pulled back.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#767

Earlier quoted context omitted.

You can't copyright AI generated works. OpenAI are barking up the wrong tree.

They're not making a legal claim, they're trying to establish provenance over Deepseek in the public eye.

Yes; and trying to justify their own valuation by pointing out that Deepseek cost more than advertised to create if you count in the cost of creating OpenAI's model.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#768
post #758

Hard to really have any sympathy for OpenAI's position when they're actively stealing content, ignoring requests to stop then spending huge amounts to get around sites running ai poisoning scripts, making it clear they'll still take your content regardless of if you consent to it.

Can someone with more expertise help me understand what I'm looking at here? https://crt.sh/?id=10106356492

It looks like Deepseek had a subdomain called "openai-us1.deepseek.com". What is a legitimate use-case for hosting an openai proxy(?) on your subdomain like this?

Not implying anything's off here, but it's interesting to me that this OpenAI entity is one of the few subdomains they have on their site

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#769
post #389

Earlier quoted context omitted.

That's not correct. First of all, training off of data generated by another AI is generally a bad idea because you'll end up with a strictly less accurate model (usually). But secondly, and more to your point, even if you were to use training data from another model, YOU STILL NEED TO DO ALL THE TRAINING. Using data from another model won't save you any training time.

> First of all, training off of data generated by another AI is generally a bad idea because you'll end up with a strictly less accurate model (usually). That is not true at all. We have known how to solve this for at least 2 years now. All the latest state of the art models depend heavily on training on synthetic data.

https://www.nature.com/articles/s41586-024-07566-y

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#770

Earlier quoted context omitted.

[flagged]

What double standard? I have not claimed anything about openAI being ethical paragon of virtue. I anything, you can be accused of using whataboutism in order to justify DeepSeek illegal actions

I'm not talking about OpenAI, I'm pointing out the unnecessary sideswipes directed at China whenever they do something bad that the US pioneered.

It's hard to read those as anything other than casual racism.

Post reply on HN