Live data from Hacker News

OpenAI says it has evidence DeepSeek used its model to train competitor

ft.com

781–790 of 1001 posts

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#781

All the top level comments are basking in the irony of it, which is fair enough. But I think this changes the Deepseek narrative a bit. If they just benefited from repurposing OpenAI data, that's different than having achieved an engineering breakthrough, which may suggest OpenAI's results were hard earned after all.

> If they just benefited from repurposing OpenAI data, that's different than having achieved an engineering breakthrough

One way or another, they were able to create something that has WAY cheaper inference costs than o1 at the same level of intelligence. I was paying Anthropic $15/1M tokens to make myself 10x faster at writing software, which was coming out to $10/day. O1 is $60/1M tokens, which for my level of usage would mean that it costs as much as a whole junior software engineer. DeepSeek is able to do it for $2.50/1M tokens.

Either OpenAI was taking a profit margin that would make the US Healthcare industry weep, or DeepSeek made an engineering breakthrough that increases inference efficiency by orders of magnitude.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#782
post #320

I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…

“We engage in countermeasures to protect our IP, including a careful process for which frontier capabilities to include in released models, and believe . . . it is critically important that we are working closely with the US government to best protect the most capable models from efforts by adversaries and competitors to take US technology.”

The above OpenAI quote from the article leans heavily towards #1 and IMO not at all towards #2. The later would be an extremely charitable reading of their statement.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#783
post #776
post #773

Earlier quoted context omitted.

Have you got some support for this claim? There's a lot of wild claims about, so while this is plausible it would be great if there were some evidence backing it.

NYT claims that OpenAI trained on their material. They argue for copyright violation, although I think another argument might be breach of TOS in scraping the material from their website or archive. The complaint filing has some references to some of the other training material used by OpenAI, but I didn't dig deeply in to what all of it was: https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20...

What's that got to do with this books claim?

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#784
post #782
post #320

I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…

“We engage in countermeasures to protect our IP, including a careful process for which frontier capabilities to include in released models, and believe . . . it is critically important that we are working closely with the US government to best protect the most capable models from efforts by adversaries and competitors to take US technology.” The above OpenAI quote from the article leans heavily towards #1 and IMO not…

What they say explicitly is not what they say implicitly. PR is an art.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#785

I think OpenAI is in a really weak position here. There are essentially two positions you can be in: You can be the agile new startup that can break the rules and move fast. That's what OpenAI used to be. Or you can be the big incumbent who is going to use your enormous resources to crush your opposition. That's Google & Microsoft here. For Microsoft to say "We're going to tie you up in lawsuits about the way you tra…

> they're still a minnow 3K+ employees, $3B+ revenue, ... sure, not BigTech but hardly a minnow. A company that big can chew gum and walk at the same time.

$7-8b in costs so they're losing $5b

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#786
I wish there were a stock ticker for OpenAI just to see what wall street’s take on all this is. One can imagine based on Nvidia, but I imagine OpenAI private valuation is hit much harder. Still, I think they’ll be able to justify it by building amazing products. Just interesting to watch what bankers think.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#788
post #320

I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…

Guess it is a good thing the AI output can’t be copyrighted, so at most they violated a policy.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#790

Earlier quoted context omitted.

Actually the "our IP" argument is ridiculous. What they are doing is stealing data from all over the web, without people's consent for that data to be used in training ML models. If anything, then "Open"AI should be sued and forced to publish their whole product. The people should demand knowing exactly what is going on with their data. Also still an unresolved issue is how they will ever comply with a deletion reque…

If there's any litigation, a counterclaim would be interesting. But DeepSeek would need to partner with parties that have been damaged by OpenAI's scraping.

I'm getting popcorn ready for the trial where an apparatus of the Chinese Communist Party files a counterclaim in an American Court together with the common people - millions of John Does - as litigants against an organization that has aggressively and in many cases of oppressively scraped their websites (DDoS)
Post reply on HN