There is a lot of discussion here about IP theft. Honest question, from deepseek's point of view as a company under a different set of laws than US/Western -- was there IP theft? A company like OpenAI can put whatever licensing they want in place. But that only matters if they can enforce it. The question is, can they enforce it against deepseek? Did deepseek do something illegal under the laws of their originating c…
The most interesting part is that China has been ahead of the US in AI for many years, just not in LLMs. You need to visit mainland China and see how AI applications are everywhere, from transport to goods shipping. I'm not surprised at all. I hope this in the end makes the US kill its strict IP laws, which is the problem. If the US doesn't, China will always have a huge edge on it, no matter how much NVidia hardware…
OpenAI says it has evidence DeepSeek used its model to train competitor
401–410 of 1001 posts
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#402Earlier quoted context omitted.
I saw a some Europeans hoping that the US would ban DeepSeek, because then there would be less traffic interfering with their own DeepSeek queries. The US can ban all they want, but if the rest of the world starts preferring Chinese social media, Chinese AI, and Chinese websites in general, the US is going to lose one of its crown jewels. The way the US behaves is a problem and makes a lot of people prefer alternativ…
EU will ban DeepSeek sooner because of (lack of) GDPR compliance
If DeepSeek becomes popular in America I predict it will be blocked, national firewall style. Will EU do the same?
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#403Earlier quoted context omitted.
That's not correct. First of all, training off of data generated by another AI is generally a bad idea because you'll end up with a strictly less accurate model (usually). But secondly, and more to your point, even if you were to use training data from another model, YOU STILL NEED TO DO ALL THE TRAINING. Using data from another model won't save you any training time.
I think the point is that if R1 isn't possible without access to OpenAI (at low, subsidized costs) then this isn't really a breakthrough as much as a hack to clone an existing model.
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#404Everyone is responding to the intellectual property issue, but isn't that the less interesting point? If Deepseek trained off OpenAI, then it wasn't trained from scratch for "pennies on the dollar" and isn't the Sputnik-like technical breakthrough that we've been hearing so much about. That's the news here. Or rather, the potential news, since we don't know if it's true yet.
That's not correct. First of all, training off of data generated by another AI is generally a bad idea because you'll end up with a strictly less accurate model (usually). But secondly, and more to your point, even if you were to use training data from another model, YOU STILL NEED TO DO ALL THE TRAINING. Using data from another model won't save you any training time.
Ah. So if I understand this... once the internet becomes completely overrun with AI-generated articles of no particular substance or importance, we should not bulk-scrape that internet again to train the subsequent generation of models.
I look forward to that day.
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#405All the top level comments are basking in the irony of it, which is fair enough. But I think this changes the Deepseek narrative a bit. If they just benefited from repurposing OpenAI data, that's different than having achieved an engineering breakthrough, which may suggest OpenAI's results were hard earned after all.
I understand they just used the API to talk to the OpenAI models. That... seems pretty innocent? Probably they even paid for it? OpenAI is selling API access, someone decided to buy it. Good for OpenAI! I understand ToS violations can lead to a ban. OpenAI is free to ban DeepSeek from using their APIs.
If your own API can leak your secret sauce without any malicious penetration, well, that's on you.
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#406Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#407We demand immediate government action to prevent these cheaper foreign AIs from taking jobs away from our great American AIs!
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#408Earlier quoted context omitted.
That's not correct. First of all, training off of data generated by another AI is generally a bad idea because you'll end up with a strictly less accurate model (usually). But secondly, and more to your point, even if you were to use training data from another model, YOU STILL NEED TO DO ALL THE TRAINING. Using data from another model won't save you any training time.
I think the point is that if R1 isn't possible without access to OpenAI (at low, subsidized costs) then this isn't really a breakthrough as much as a hack to clone an existing model.
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#409"You can't take data without asking" seems like a court precedent OpenAI really, really, really wants to avoid. And yet...
OpenAI is saying that their service was used in violation of their TOS, which is a bit different than just copying data. To be clear I’m not on OpenAI’s side, but it looks to me that the legal situation isn’t exactly analogous.
Which is the most ridiculous argument they could use because they didn't respect any ToS (or copyright laws, for that matter) when scraping the whole web, books from Libgen and who knows what more.
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#410Everyone is responding to the intellectual property issue, but isn't that the less interesting point? If Deepseek trained off OpenAI, then it wasn't trained from scratch for "pennies on the dollar" and isn't the Sputnik-like technical breakthrough that we've been hearing so much about. That's the news here. Or rather, the potential news, since we don't know if it's true yet.
That's not correct. First of all, training off of data generated by another AI is generally a bad idea because you'll end up with a strictly less accurate model (usually). But secondly, and more to your point, even if you were to use training data from another model, YOU STILL NEED TO DO ALL THE TRAINING. Using data from another model won't save you any training time.
It proofs we _can_ optimize our training data.
Just like humans have been genetically stable for a long time, the quality & structure of information available to a child today vs that of 2000 years ago makes them more skilled at certain tasks. Math being a good example.