Live data from Hacker News

OpenAI says it has evidence DeepSeek used its model to train competitor

ft.com

211–220 of 1001 posts

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#211

I do think that distilling a model from another is much less impressive than distilling one from raw text. However, it is hard to say if it is really illegal or even immoral, perhaps just one step further in the evolution of the space.

It's about as illegal as the billions, if not trillions of IPs that ClosedAI infringed to train their own data without consent. Not that they're alone, and I personally don't mind that AI companies do it, but it's still amusing when they get this annoyed at others doing the same thing to them.

I think they had the advantage of being ahead of the law in this regard. To my knowledge, reading copywritten material isn't (or wasn't illegal) and remains a legal grey area.

Distilling weights from prompts and responses is even more of a legal grey area. The legal system cannot respond quickly to such technological advancements so things necessarily remain a wild west until technology reaches the asymptotic portion of the curve.

In my view the most interesting thing is, do we really need vast data centers and innumerable GPUs for AGI? In other words, if intelligence is ultimately a function of power input, what is the shape of the curve?

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#212

It's reasonably likely that a lot of people linked to the federal government want to ban DeepSeek. You can tell it's being presented away from "they gave us a free set of weights" and towards "they destroyed $1T of shareholder value." (By revealing that Microsoft et al. paid way too much to OpenAI et al. for technology that was actually easy to reinvent.)

I saw a some Europeans hoping that the US would ban DeepSeek, because then there would be less traffic interfering with their own DeepSeek queries. The US can ban all they want, but if the rest of the world starts preferring Chinese social media, Chinese AI, and Chinese websites in general, the US is going to lose one of its crown jewels. The way the US behaves is a problem and makes a lot of people prefer alternativ…

Agreed, you've highlighted one of the key problems with protectionism and nativism. Banning competition just weakens America's global influence, it doesn't make it stronger.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#213
post #208

Earlier quoted context omitted.

[flagged]

Yeah I don't know, Altman is a sociopath who is now trying to get intertwined with local governments (SF) as well as the federal government. He's going to do a lot of weaseling to get what he wants: laws that forcibly make OpenAI a monopoly. Society will always have crazy sociopaths destroying things for their own gain, and now is Altman's turn.

[flagged]

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#214

It's reasonably likely that a lot of people linked to the federal government want to ban DeepSeek. You can tell it's being presented away from "they gave us a free set of weights" and towards "they destroyed $1T of shareholder value." (By revealing that Microsoft et al. paid way too much to OpenAI et al. for technology that was actually easy to reinvent.)

Not sure I agree with your premise, but what exactly are they going to ban?

They can stop DeepSeek doing various things commercially I guess, but stopping Americans using their ideas is simply impossible and stopping use of their source or weights would be (likely successfully) challenged under the first amendment.

There is no law against simply destroying trillions of dollars of shareholder value.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#215
post #21

This is absolutely hilarious! :) ClosedAI scraped human content without asking and they explained why this was acceptable... but when the outputs of their training corpus is scraped, it is THEIR dataset and this is NOT acceptable! Oh, the irony! :D I shared a few screenshots of DeepSeek answering using ChatGPT's output in yesterday's article! https://semking.com/deepseek-china-ai-model-breakthrough-sec...

Does this mean when you use OpenAI as an enterprise customer, they can see exactly the queries and answers? So much for privacy!

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#216

Earlier quoted context omitted.

[flagged]

I personally love this chef's kiss of a flip flop sam did here: https://blog.samaltman.com/trump https://www.reddit.com/r/YAPms/comments/1i7ry5m/sam_altman_g... Only a truly talented piece of shit can be as prolific as this. "He is irresponsible in the way dictators are." Chef's kiss. Edit: Kids, don't aspire to be like Altman. We as a community need to espouse more values than tech is gonna tech .

I especially like how he quoted Napoleon or something framing himself as the heart of revolution and Deep Seek as a child of the revolution only to get a response from some random guy "It's not that deep bro. Just release a better model."

https://x.com/hibakod/status/1883189126553596234

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#217

OpenAI is also possibly in violation of many IP laws by scraping the entirety of the internet and using to train their models, so there’s that.

To my understanding, OpenAI won the case where it argued training was covered under fair use and did not infringe on copyright.

Is there any reason they wouldn't rule the same way on DeepSeek training on OpenAI data? After all, one of the big selling points of GPT has been that businesses can freely use the information provided. They're paying for the service, after all. I'd very be interested to know how DeepSeek's usage (very reasonably assuming that they paid for their OpenAI subscription) is any different.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#218
post #192

"You can't take data without asking" seems like a court precedent OpenAI really, really, really wants to avoid. And yet...

OpenAI is saying that their service was used in violation of their TOS, which is a bit different than just copying data. To be clear I’m not on OpenAI’s side, but it looks to me that the legal situation isn’t exactly analogous.

Tons of websites and books they scraped had copyright notices.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#219

Earlier quoted context omitted.

[flagged]

I personally love this chef's kiss of a flip flop sam did here: https://blog.samaltman.com/trump https://www.reddit.com/r/YAPms/comments/1i7ry5m/sam_altman_g... Only a truly talented piece of shit can be as prolific as this. "He is irresponsible in the way dictators are." Chef's kiss. Edit: Kids, don't aspire to be like Altman. We as a community need to espouse more values than tech is gonna tech .

That is particularly gross, but that really feels like the norm among all the tech elite these days - Zuckerberg, Bezos, etc. all doing the most laughable flip flops.

The reason the flip flops are so laughable to me is because they attempt to couch them in some noble, moralistic viewpoint, instead of the obvious reason "We own big companies, the government has extreme power to make or break these companies, and everyone knows kissing up to Trump is what is required to be on his good side."

Profiles in Cowardice, every last one of them.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#220
All the top level comments are basking in the irony of it, which is fair enough. But I think this changes the Deepseek narrative a bit. If they just benefited from repurposing OpenAI data, that's different than having achieved an engineering breakthrough, which may suggest OpenAI's results were hard earned after all.
Post reply on HN