Live data from Hacker News

OpenAI says it has evidence DeepSeek used its model to train competitor

ft.com

131–140 of 1001 posts

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#133

It's reasonably likely that a lot of people linked to the federal government want to ban DeepSeek. You can tell it's being presented away from "they gave us a free set of weights" and towards "they destroyed $1T of shareholder value." (By revealing that Microsoft et al. paid way too much to OpenAI et al. for technology that was actually easy to reinvent.)

"easy to reinvent" often comes after "hard to invent"

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#135
post #21

This is absolutely hilarious! :) ClosedAI scraped human content without asking and they explained why this was acceptable... but when the outputs of their training corpus is scraped, it is THEIR dataset and this is NOT acceptable! Oh, the irony! :D I shared a few screenshots of DeepSeek answering using ChatGPT's output in yesterday's article! https://semking.com/deepseek-china-ai-model-breakthrough-sec...

openai should pay creators, but: 1. scraping the internet and making AI out of it 2. using the AI from #1 to create another AI are not the same thing.

Yes, they are different actions.

But arguably these actions share enough characteristics that it’s reasonable to place them in the same category. Something like: “products that exist largely/solely because of the work of other people”. The nonconsensual nature of this and the lack of compensation is what people understandably take issue with.

There is enough similarity that it evokes specific feelings about OpenAI when they suddenly find themselves on the other side of the situation.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#136
post #126
post #21

This is absolutely hilarious! :) ClosedAI scraped human content without asking and they explained why this was acceptable... but when the outputs of their training corpus is scraped, it is THEIR dataset and this is NOT acceptable! Oh, the irony! :D I shared a few screenshots of DeepSeek answering using ChatGPT's output in yesterday's article! https://semking.com/deepseek-china-ai-model-breakthrough-sec...

While all of this is true, that DeepSeek wouldn't be here were it not for the research that preceded it notably Google's paper, then Llama, and ChatGPT which they're modeled after, its release still did something profound to their psyche, the motivation and self-actualization this instills to the Chinese. They witnessed the power of their accomplishments: a side-hustle project knocked off an easy trillion. This is on…

OpenAI wouldn't be here without the work that Yann Lecun did at Facebook (back when it was facebook). Science is built on top of science, that's just how things work.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#137

Earlier quoted context omitted.

Also, DeepSeek is allegedly... better? So saying they just copied ClosedAI isn't really sufficient of an answer. Seems to be just bluster because the US Govt would probably accept any excuse to ban it, see TikTok.

It’s not better. In most of my tests (C++/QT code) it just runs out of context before it can really do anything. And the output is very bad - it mashes together the header and cpp file. The reasoning output is fun to look at and occasionally useful though. The max token output is only 8K (32K thinking tokens). O1 is 128k, which is far more useful, and it doesn’t get stuck like R1 does. The hype around the DeepSeek re…

Is this a local run of one of the smaller models and/or other-models-distilled-with-r1, or are you using their Chat interface?

I've also compared o1 and (online-hosted) r1 on Qt/C++ code, being a KDE Plasma dev, and my impression so far was that the output is roughly on par. I've given both models some tricky tasks about dark corners of the meta-object system in crafting classes etc. and they came up with generally the same sort of suggestions and implementations.

I do appreciate that "asking about gotchas with few definitive solutions, even if they require some perspective" and "rote day-to-day coding ops" are very different benchmarks due to how things are represented in the training data corpus, though.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#138
post #21

This is absolutely hilarious! :) ClosedAI scraped human content without asking and they explained why this was acceptable... but when the outputs of their training corpus is scraped, it is THEIR dataset and this is NOT acceptable! Oh, the irony! :D I shared a few screenshots of DeepSeek answering using ChatGPT's output in yesterday's article! https://semking.com/deepseek-china-ai-model-breakthrough-sec...

openai should pay creators, but: 1. scraping the internet and making AI out of it 2. using the AI from #1 to create another AI are not the same thing.

I'm genuinely not sure which one you think is worse (if any). (1) seems worse, but your reply suggests to me maybe you think (2) is worse.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#140

It's reasonably likely that a lot of people linked to the federal government want to ban DeepSeek. You can tell it's being presented away from "they gave us a free set of weights" and towards "they destroyed $1T of shareholder value." (By revealing that Microsoft et al. paid way too much to OpenAI et al. for technology that was actually easy to reinvent.)

Theoretically this should be good for OpenAI - in that they can reduce their costs by ~27x and pass that along to end users to get more adoption and more profit.

Nah, those costs were for their doomsday bunkers and crypto purchases, and maybe a house or 3
Post reply on HN