Live data from Hacker News

OpenAI says it has evidence DeepSeek used its model to train competitor

ft.com

391–400 of 1001 posts

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#392
post #374

Everyone is responding to the intellectual property issue, but isn't that the less interesting point? If Deepseek trained off OpenAI, then it wasn't trained from scratch for "pennies on the dollar" and isn't the Sputnik-like technical breakthrough that we've been hearing so much about. That's the news here. Or rather, the potential news, since we don't know if it's true yet.

This has been in the back of my head since the news broke. Has anyone built their own R1 from scratch and validated it?

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#393
post #257

It's reasonably likely that a lot of people linked to the federal government want to ban DeepSeek. You can tell it's being presented away from "they gave us a free set of weights" and towards "they destroyed $1T of shareholder value." (By revealing that Microsoft et al. paid way too much to OpenAI et al. for technology that was actually easy to reinvent.)

There's a lot of egg on people's faces now. DeepSeek shows there's nothing special about America or its economic system that breeds innovation. DeepSeek shows how these tech oligarchs greatly overplayed their hand and along with the president, bamboozled the taxpayer to enrich each other. I just hope the voters remember this in 2 years, 4 years and beyond.

If the allegations are true, the special thing about OpenAI is that it didn't have to be trained off DeepSeek. But either way, you maybe don't want to invest billions in something if someone else will be able to copy it for less.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#394
post #370

Earlier quoted context omitted.

[flagged]

Ok, but please don't break HN's rules when commenting here. You may not owe altmen better, but you owe this community better if you're participating in it. https://news.ycombinator.com/newsguidelines.html

Once again you abuse your moderator powers to enforce your personal vendetta against people who dare to speak ill of tech CEOs.

I find your behavior repulsive and fervently wish you would quit.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#395
post #389
post #374

Everyone is responding to the intellectual property issue, but isn't that the less interesting point? If Deepseek trained off OpenAI, then it wasn't trained from scratch for "pennies on the dollar" and isn't the Sputnik-like technical breakthrough that we've been hearing so much about. That's the news here. Or rather, the potential news, since we don't know if it's true yet.

That's not correct. First of all, training off of data generated by another AI is generally a bad idea because you'll end up with a strictly less accurate model (usually). But secondly, and more to your point, even if you were to use training data from another model, YOU STILL NEED TO DO ALL THE TRAINING. Using data from another model won't save you any training time.

I think the point is that if R1 isn't possible without access to OpenAI (at low, subsidized costs) then this isn't really a breakthrough as much as a hack to clone an existing model.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#396
post #333

Earlier quoted context omitted.

Even more hilarious given their own charter: > We will attempt to directly build safe and beneficial AGI, but will also consider our mission fulfilled if our work aids others to achieve this outcome. > Our primary fiduciary duty is to humanity. We anticipate needing to marshal substantial resources to fulfill our mission, but will always diligently act to minimize conflicts of interest among our employees and stakeho…

Ah yes: "duty to humanity"

I think one good thing to come out of all this tech elite flip flopping is that I now see these tech leaders for exactly who they are. It makes me kind of sad, because as someone who came of age early in the Web era I really wanted to believe that there was a bigger moral good to all we were doing.

I now view any moralistic statement by any of these big tech companies as complete and total bullshit, which is probably for the best, because that is what it is. These companies now exist solely to amass power and wealth. They will still use moralistic language to try to motivate their employees, but I hope folks still see it for the complete nonsense that it is.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#397
post #60

There is an Egyptian say that would translate to something like "We didn’t see them when they were stealing, we saw them when they were fighting over what was stolen" That describes this situation. Although to be honest all this aggressive scraping is noticeable but for people who understand that which is not majority of people. but now everyone knows.

> Although to be honest all this aggressive scraping is noticeable but for people who understand that which is not majority of people.

When you say noticeable, do you mean in like, traffic statistics? Or in what the model knows that it clearly shouldn't if it wasn't trained in legally dubious ways?

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#398

It's reasonably likely that a lot of people linked to the federal government want to ban DeepSeek. You can tell it's being presented away from "they gave us a free set of weights" and towards "they destroyed $1T of shareholder value." (By revealing that Microsoft et al. paid way too much to OpenAI et al. for technology that was actually easy to reinvent.)

I saw a some Europeans hoping that the US would ban DeepSeek, because then there would be less traffic interfering with their own DeepSeek queries. The US can ban all they want, but if the rest of the world starts preferring Chinese social media, Chinese AI, and Chinese websites in general, the US is going to lose one of its crown jewels. The way the US behaves is a problem and makes a lot of people prefer alternativ…

I've recently cancelled my Github Copilot subscription and now use Mistral. When the US starts threatening allies with tariffs or invasion, using US services becomes a major business risk.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#399
post #224
post #192

Earlier quoted context omitted.

OpenAI is saying that their service was used in violation of their TOS, which is a bit different than just copying data. To be clear I’m not on OpenAI’s side, but it looks to me that the legal situation isn’t exactly analogous.

If using data violating some ToS taints the model trained on that data, then all of OpenAI's models are tainted by the millions of ToS'es they broke.

Can you cite a source showing they violated ToS?

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#400
post #389
post #374

Everyone is responding to the intellectual property issue, but isn't that the less interesting point? If Deepseek trained off OpenAI, then it wasn't trained from scratch for "pennies on the dollar" and isn't the Sputnik-like technical breakthrough that we've been hearing so much about. That's the news here. Or rather, the potential news, since we don't know if it's true yet.

That's not correct. First of all, training off of data generated by another AI is generally a bad idea because you'll end up with a strictly less accurate model (usually). But secondly, and more to your point, even if you were to use training data from another model, YOU STILL NEED TO DO ALL THE TRAINING. Using data from another model won't save you any training time.

I think you're missing the point being made here, IMHO: using an advanced model to build high quality training data (whatever that means for a given training paradigm) absolutely would increase the efficiency of the process. Remember that they're not fighting over sounding human, they're fighting over deliberative reasoning capabilities, something that's relatively rare in online discourse.

Re: "generally a bad idea", I'd just highlight "generally" ;) Clearly it worked in this case!

Post reply on HN