Live data from Hacker News

OpenAI says it has evidence DeepSeek used its model to train competitor

ft.com

251–260 of 1001 posts

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#251

Earlier quoted context omitted.

[flagged]

I personally love this chef's kiss of a flip flop sam did here: https://blog.samaltman.com/trump https://www.reddit.com/r/YAPms/comments/1i7ry5m/sam_altman_g... Only a truly talented piece of shit can be as prolific as this. "He is irresponsible in the way dictators are." Chef's kiss. Edit: Kids, don't aspire to be like Altman. We as a community need to espouse more values than tech is gonna tech .

I don’t think your links are evidence of a flip flop.

The first link is from mid-2016. The second link is from January 2025.

It is entirely reasonable for someone to genuinely change his or her views of a person over the course of 8.5 years. That is a substantial length of time in a person’s life.

To me a “flip-flop” is when one changes views on something in a very short amount of time.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#252

Earlier quoted context omitted.

Also, DeepSeek is allegedly... better? So saying they just copied ClosedAI isn't really sufficient of an answer. Seems to be just bluster because the US Govt would probably accept any excuse to ban it, see TikTok.

It’s not better. In most of my tests (C++/QT code) it just runs out of context before it can really do anything. And the output is very bad - it mashes together the header and cpp file. The reasoning output is fun to look at and occasionally useful though. The max token output is only 8K (32K thinking tokens). O1 is 128k, which is far more useful, and it doesn’t get stuck like R1 does. The hype around the DeepSeek re…

It's not great at super-complex tasks due to limited context, but it's quite a good "junior intern that has memorized the Internet." Local deepseek-r1 on my laptop (M1 w/64GiB RAM) can answer about any question I can throw at it... as long as it's not something on China's censored list. :)

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#254
post #21

This is absolutely hilarious! :) ClosedAI scraped human content without asking and they explained why this was acceptable... but when the outputs of their training corpus is scraped, it is THEIR dataset and this is NOT acceptable! Oh, the irony! :D I shared a few screenshots of DeepSeek answering using ChatGPT's output in yesterday's article! https://semking.com/deepseek-china-ai-model-breakthrough-sec...

I think the point is - OpenAI scraped public data - d1 - Trained their model to produce output - d2 - DeepSeek used d2 to reinforce their model

OpenAI is mad about d2 (not d1). I'm not sure using public data is "stealing". In summary, these are two different things & need to be separate.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#255
post #21

This is absolutely hilarious! :) ClosedAI scraped human content without asking and they explained why this was acceptable... but when the outputs of their training corpus is scraped, it is THEIR dataset and this is NOT acceptable! Oh, the irony! :D I shared a few screenshots of DeepSeek answering using ChatGPT's output in yesterday's article! https://semking.com/deepseek-china-ai-model-breakthrough-sec...

[flagged]

I don’t care for Sam Altman and his general untrustworthy behavior. But DeepSeek is perhaps more untrustworthy. Models from American companies at least aren’t surprising us with government driven misinformation, and even though safety can also be censorship, the companies that make these models at least openly talk about their safety programs. DeepSeek is implementing a censorship and propaganda program without admitting it at all, and once they become good at doing it in less obvious ways, it can become very damaging and corrupt the political process of other societies, because users will trust the tools they use are neutral.

I think DeepSeek’s strategy to announce a misleading low cost (just the final training run that optimizes a base model that in turn is possibly based on OpenAI) is also purposeful. After all, High Flyer, the parent company of DeepSeek, is a hedge fund - and I bet they took out big short positions on Nvidia before their recent announcements. The Chinese government, of course, benefits from a misleading number being announced broadly, causing doubt among investors who would otherwise continue to prop up American technology startups. Not to mention the big fall in American markets as a result.

I do think there’s also a big difference between scraping the Internet for training data, which might just be fair use, and training off other LLMs or obtaining their assets in some other way. The latter feels like the kind of copying and industrial espionage that used to get China ridiculed in the 2000s and 2010s. Note that DeepSeek has never detailed their training data, even at a high level. This is true even in their previous papers, where they were very vague about the pre training process, which feels suspicious.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#256

Earlier quoted context omitted.

>The US government likely will favor a large strategic company like OpenAI instead of individual's copyrights Even if we assume this is true, Disney and Netflix are both currently worth more than OpenAI and both rely on the strict enforcement of US copyright law. I do not think it is so obvious which powers that be have the better lobbying efforts and, currently, it's looking like this question will mostly be adjudic…

I don't think OpenAI stole from Disney or Netflix. Rather OpenAI stole from individual artists and YouTube and other social media who users do not really have any lobbying power. So I think OpenAI, Disney and Netflix win together. Big companies tend to win.

Disney owns ABC News; OpenAI almost certainly scraped their text data

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#257

It's reasonably likely that a lot of people linked to the federal government want to ban DeepSeek. You can tell it's being presented away from "they gave us a free set of weights" and towards "they destroyed $1T of shareholder value." (By revealing that Microsoft et al. paid way too much to OpenAI et al. for technology that was actually easy to reinvent.)

There's a lot of egg on people's faces now. DeepSeek shows there's nothing special about America or its economic system that breeds innovation. DeepSeek shows how these tech oligarchs greatly overplayed their hand and along with the president, bamboozled the taxpayer to enrich each other. I just hope the voters remember this in 2 years, 4 years and beyond.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#258

Earlier quoted context omitted.

For the curious, it was vertical integration in the railroad-oil/-coal industry which is where the money was made. The problem for AI is the hardware is commodified and offers no natural monopoly, so there isn't really anything obvious to vertically integrate-towards-monopoly.

Aren’t we approaching a scenario where the software is commodified (or at least “good enough” software) and the hardware isn’t (NVIDIA GPUs have defined advantages)

I think the lesson of DeepSeek is 'no' -- that by software innovation (ie., dropping below CUDA to programming the GPU directly, working at 8bit, etc.) you can trivialise the hardware requirement.

However I think the reality is that there's only so much coal to be mined, as far as LLM training goes. When we're at "very dimishing returns" SoC/Apple/TSMC-CPU innovations will deliver cheap inference. We only really need a M4 Ultra with 1TB RAM to hollow-out the hardware-inference-supplier market.

Very easy to imagine a future where Apple releases a "Apple Intelligence Mac Studio" with the specs for many businesses to run arbitrary models.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#259

Earlier quoted context omitted.

> "they destroyed $1T of shareholder value." (By revealing that Microsoft et al. paid way too much to OpenAI et al. for technology that was actually easy to reinvent.) The value was highly speculative, an illusion created by PR and sentiment momentum. "Hype value" not real value (unless you're able to realize it and dump those bags on someone else before fundamentals set in). Same thing happening with power companies…

The two podcasters who do the Acquired podcast spoke to Ballmer about some of Microsoft’s failed initiatives and acquisitions. He told them that at the end of the day “it’s only money”. All of the BigTech companies have enough cash flow from profitable lines of business to make speculative bets.

It must be EZ mode to be a big tech executive, you somehow have all the power to make every decision while also having the ability to never take the fault for these decisions.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#260
post #192

Earlier quoted context omitted.

OpenAI is saying that their service was used in violation of their TOS, which is a bit different than just copying data. To be clear I’m not on OpenAI’s side, but it looks to me that the legal situation isn’t exactly analogous.

Tons of websites and books they scraped had copyright notices.

Copyright and terms of service are different legal notions.
Post reply on HN