Live data from Hacker News

OpenAI says it has evidence DeepSeek used its model to train competitor

ft.com

401–410 of 1001 posts

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#401

There is a lot of discussion here about IP theft. Honest question, from deepseek's point of view as a company under a different set of laws than US/Western -- was there IP theft? A company like OpenAI can put whatever licensing they want in place. But that only matters if they can enforce it. The question is, can they enforce it against deepseek? Did deepseek do something illegal under the laws of their originating c…

The most interesting part is that China has been ahead of the US in AI for many years, just not in LLMs. You need to visit mainland China and see how AI applications are everywhere, from transport to goods shipping. I'm not surprised at all. I hope this in the end makes the US kill its strict IP laws, which is the problem. If the US doesn't, China will always have a huge edge on it, no matter how much NVidia hardware…

You don’t even need to visit china, just read the latest research papers and look at the authors. China has more researchers in AI than the West and that’s a proven way to build an advantage.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#402

Earlier quoted context omitted.

I saw a some Europeans hoping that the US would ban DeepSeek, because then there would be less traffic interfering with their own DeepSeek queries. The US can ban all they want, but if the rest of the world starts preferring Chinese social media, Chinese AI, and Chinese websites in general, the US is going to lose one of its crown jewels. The way the US behaves is a problem and makes a lot of people prefer alternativ…

EU will ban DeepSeek sooner because of (lack of) GDPR compliance

Does EU block websites that don't comply with their laws?

If DeepSeek becomes popular in America I predict it will be blocked, national firewall style. Will EU do the same?

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#403
post #389

Earlier quoted context omitted.

That's not correct. First of all, training off of data generated by another AI is generally a bad idea because you'll end up with a strictly less accurate model (usually). But secondly, and more to your point, even if you were to use training data from another model, YOU STILL NEED TO DO ALL THE TRAINING. Using data from another model won't save you any training time.

I think the point is that if R1 isn't possible without access to OpenAI (at low, subsidized costs) then this isn't really a breakthrough as much as a hack to clone an existing model.

The training techniques are a breakthrough no matter what data is used. It's not up for debate, it's an empirical question with a concrete answer. They can and did train orders of magnitude faster.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#404
post #389
post #374

Everyone is responding to the intellectual property issue, but isn't that the less interesting point? If Deepseek trained off OpenAI, then it wasn't trained from scratch for "pennies on the dollar" and isn't the Sputnik-like technical breakthrough that we've been hearing so much about. That's the news here. Or rather, the potential news, since we don't know if it's true yet.

That's not correct. First of all, training off of data generated by another AI is generally a bad idea because you'll end up with a strictly less accurate model (usually). But secondly, and more to your point, even if you were to use training data from another model, YOU STILL NEED TO DO ALL THE TRAINING. Using data from another model won't save you any training time.

> training off of data generated by another AI is generally a bad idea

Ah. So if I understand this... once the internet becomes completely overrun with AI-generated articles of no particular substance or importance, we should not bulk-scrape that internet again to train the subsequent generation of models.

I look forward to that day.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#405
post #364

All the top level comments are basking in the irony of it, which is fair enough. But I think this changes the Deepseek narrative a bit. If they just benefited from repurposing OpenAI data, that's different than having achieved an engineering breakthrough, which may suggest OpenAI's results were hard earned after all.

I understand they just used the API to talk to the OpenAI models. That... seems pretty innocent? Probably they even paid for it? OpenAI is selling API access, someone decided to buy it. Good for OpenAI! I understand ToS violations can lead to a ban. OpenAI is free to ban DeepSeek from using their APIs.

That's how I understand it too.

If your own API can leak your secret sauce without any malicious penetration, well, that's on you.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#406
It’s not a good look when your technology is replicated for a fraction of the cost, and your response is to smear your competition with (probably) false accusations and cozy up to the US government to tighten already shortsighted export controls. Hubris & xenophobia are not going to serve American companies well. Personally I welcome the Chinese - or anyone else for that matter - developing advanced technologies as long as they are used for good. Humanity loses if we allow this stuff to be “owned” by a handful of companies or a single country.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#407

We demand immediate government action to prevent these cheaper foreign AIs from taking jobs away from our great American AIs!

does it matter if the company gets banned? other non-chinese companies can pick up the open source model and run it as a service with relatively low investment, isn't that the point?

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#408
post #389

Earlier quoted context omitted.

That's not correct. First of all, training off of data generated by another AI is generally a bad idea because you'll end up with a strictly less accurate model (usually). But secondly, and more to your point, even if you were to use training data from another model, YOU STILL NEED TO DO ALL THE TRAINING. Using data from another model won't save you any training time.

I think the point is that if R1 isn't possible without access to OpenAI (at low, subsidized costs) then this isn't really a breakthrough as much as a hack to clone an existing model.

R1 is--as far as we know from good ol' ClosedAI--far more efficient. Even if it were a "clone", A) that would be a terribly impressive achievement on its own that Anthropic and Google would be mighty jealous of, and B) it's at the very least a distillation of O1's reasoning capabilities into a more svelte form.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#409
post #192

"You can't take data without asking" seems like a court precedent OpenAI really, really, really wants to avoid. And yet...

OpenAI is saying that their service was used in violation of their TOS, which is a bit different than just copying data. To be clear I’m not on OpenAI’s side, but it looks to me that the legal situation isn’t exactly analogous.

> OpenAI is saying that their service was used in violation of their TOS

Which is the most ridiculous argument they could use because they didn't respect any ToS (or copyright laws, for that matter) when scraping the whole web, books from Libgen and who knows what more.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#410
post #389
post #374

Everyone is responding to the intellectual property issue, but isn't that the less interesting point? If Deepseek trained off OpenAI, then it wasn't trained from scratch for "pennies on the dollar" and isn't the Sputnik-like technical breakthrough that we've been hearing so much about. That's the news here. Or rather, the potential news, since we don't know if it's true yet.

That's not correct. First of all, training off of data generated by another AI is generally a bad idea because you'll end up with a strictly less accurate model (usually). But secondly, and more to your point, even if you were to use training data from another model, YOU STILL NEED TO DO ALL THE TRAINING. Using data from another model won't save you any training time.

Thats not right either.

It proofs we _can_ optimize our training data.

Just like humans have been genetically stable for a long time, the quality & structure of information available to a child today vs that of 2000 years ago makes them more skilled at certain tasks. Math being a good example.

Post reply on HN