Live data from Hacker News

OpenAI says it has evidence DeepSeek used its model to train competitor

ft.com

311–320 of 1001 posts

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#311
post #265
post #95

> “It’s also extremely hard to rally a big talented research team to charge a new hill in the fog together,” he added. “This is the key to driving progress forward.” Well I think DeepSeek releasing it open source and on an MIT license will rally the big talent. The open sourcing of a new technology has always driven progress in the past. The last paragraph too is where OpenAi seems to be focusing their efforts.. > we…

or sold to US I could totally see this happening soon

Why would they want to sell ?

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#312
post #74
post #21

This is absolutely hilarious! :) ClosedAI scraped human content without asking and they explained why this was acceptable... but when the outputs of their training corpus is scraped, it is THEIR dataset and this is NOT acceptable! Oh, the irony! :D I shared a few screenshots of DeepSeek answering using ChatGPT's output in yesterday's article! https://semking.com/deepseek-china-ai-model-breakthrough-sec...

Any ML based service with an API is basically a dataset builder for more ML. This has been known forever and is actually a useful "law" of ML-based systems.

This is why models should be open. Or at least they should have a local option.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#313

We demand immediate government action to prevent these cheaper foreign AIs from taking jobs away from our great American AIs!

To the detriment of OpenAI, the math is going to be used to improve AIs developed in America. And we need to remember that Marc Andreessen is very against government banning maths.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#314

It's reasonably likely that a lot of people linked to the federal government want to ban DeepSeek. You can tell it's being presented away from "they gave us a free set of weights" and towards "they destroyed $1T of shareholder value." (By revealing that Microsoft et al. paid way too much to OpenAI et al. for technology that was actually easy to reinvent.)

I think the real concern from the govt's perspective is data privacy, since all the chat messages are stored on Chinese servers

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#315
I mean, almost ALL opensource models, ever since alpaca, contain a ton of synthetic data produced via ChatGPT in their finetuning or training datasets. It's not a surprise to anyone who's been using OSS LLMs for a while: almost ALL of them hallucinate that they are ChatGPT.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#316
post #60

There is an Egyptian say that would translate to something like "We didn’t see them when they were stealing, we saw them when they were fighting over what was stolen" That describes this situation. Although to be honest all this aggressive scraping is noticeable but for people who understand that which is not majority of people. but now everyone knows.

"We didn’t see them when we were stealing, we saw them when they were fighting over what we stole" fixed for you

That means a different thing.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#317

It's reasonably likely that a lot of people linked to the federal government want to ban DeepSeek. You can tell it's being presented away from "they gave us a free set of weights" and towards "they destroyed $1T of shareholder value." (By revealing that Microsoft et al. paid way too much to OpenAI et al. for technology that was actually easy to reinvent.)

I saw a some Europeans hoping that the US would ban DeepSeek, because then there would be less traffic interfering with their own DeepSeek queries. The US can ban all they want, but if the rest of the world starts preferring Chinese social media, Chinese AI, and Chinese websites in general, the US is going to lose one of its crown jewels. The way the US behaves is a problem and makes a lot of people prefer alternativ…

EU will ban DeepSeek sooner because of (lack of) GDPR compliance

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#318
"Stole" - I don't believe that word means what he thinks it means. Perhaps I pre-maturely anthropomorphize AI -- yet when I read a novel, such as The Sorcerer's Stone, I am not guilty of stealing Rowling's work, even if I didn't purchase the book but instead found it and read it in a friend's bathroom. Now if I were to take the specific plot and characters of that story and write a screenplay or novel directly based on it, and, explicitly, attempt to sell this work, perhaps the verb chosen here would be appropriate.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#319

It's reasonably likely that a lot of people linked to the federal government want to ban DeepSeek. You can tell it's being presented away from "they gave us a free set of weights" and towards "they destroyed $1T of shareholder value." (By revealing that Microsoft et al. paid way too much to OpenAI et al. for technology that was actually easy to reinvent.)

"easy to reinvent" often comes after "hard to invent"

At Microsoft’s size, they don’t care; they just buy out others.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#320
I think there's two different things going on here:

"DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet.

"DeepSeek trained on our outputs, and so their claims of replicating o1-level performance from scratch are not really true" This is at least plausibly a valid claim. The DeepSeek R1 paper shows that distillation is really powerful (e.g. they show Llama models get a huge boost by finetuning on R1 outputs), and if it were the case that DeepSeek were using a bunch of o1 outputs to train their model, that would legitimately cast doubt on the narrative of training efficiency. But that's a separate question from whether it's somehow unethical to use OpenAI's data the same way OpenAI uses everyone else's data.

Post reply on HN