Live data from Hacker News

OpenAI says it has evidence DeepSeek used its model to train competitor

ft.com

421–430 of 1001 posts

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#421

Earlier quoted context omitted.

The railroads drama ended when JP Morgan (the person, not yet the entity) brought all the railroad bosses together, said "you all answer to me because I represent your investors / shareholders", and forced a wave of consolidation and syndicates because competition was bad for business. Then all the farmers in the midwest went broke not because they couldn't get their goods to market, but because JP Morgan's consolida…

> Consolidation and monopoly over your competition is always the end goal. Surely that's only possible when you have a large barrier to entry? What's going to be that barrier in this case - cos it turns out not to be neither training costs/hardware or secret expertise.

So I'm not an expert in this but even with DeepSeek supposedly reducing training costs isn't the estimate still in the millions (and that's presumably not counting a lot of costs)? And that wouldn't be counting a bunch of other barriers for actually building the business since training a model is only one part, the barrier to entry still seems very high.

Also barriers to entry aren't the only way to get a consolidated market anyway.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#423

Earlier quoted context omitted.

The railroads drama ended when JP Morgan (the person, not yet the entity) brought all the railroad bosses together, said "you all answer to me because I represent your investors / shareholders", and forced a wave of consolidation and syndicates because competition was bad for business. Then all the farmers in the midwest went broke not because they couldn't get their goods to market, but because JP Morgan's consolida…

> Consolidation and monopoly over your competition is always the end goal. Surely that's only possible when you have a large barrier to entry? What's going to be that barrier in this case - cos it turns out not to be neither training costs/hardware or secret expertise.

You figure that out and the VC's will be shovelling money into your face.

I suspect the "it ain't training costs/hardware" bit is a bit exagerated since it ignores all the prior work that DeepSeek was built on top of.

But, if all else fails, there's always the tried-and-true approaches: regulatory capture, industry entrenchment, use your VC bucks to be the last one who can wait out the costs the incumbents do face before they fold, etc.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#424
post #389
post #374

Everyone is responding to the intellectual property issue, but isn't that the less interesting point? If Deepseek trained off OpenAI, then it wasn't trained from scratch for "pennies on the dollar" and isn't the Sputnik-like technical breakthrough that we've been hearing so much about. That's the news here. Or rather, the potential news, since we don't know if it's true yet.

That's not correct. First of all, training off of data generated by another AI is generally a bad idea because you'll end up with a strictly less accurate model (usually). But secondly, and more to your point, even if you were to use training data from another model, YOU STILL NEED TO DO ALL THE TRAINING. Using data from another model won't save you any training time.

> training off of data generated by another AI is generally a bad idea

It's...not, and its repeatedly been proven in practice that this is an invalid generalization because it is missing necessary qualifications, and its funny that this myth keeps persisting.

It's probably a bad idea to use uncurated output from another AI to train a model if you are trying to make a better model rather than a distillation of the first model, and its definitely (and, ISTR, the actual research result from which the false generalization has developed) a bad idea to iteratively fine-tune a model on its own unfiltered output, but there has been lots of success using AI models to generate data which is curated and used to train other models, which can be much more efficient that trying to create new material without AI once you've gotten to the point where you've already hoovered up all the readily-accessible low hanging fruit of premade content relevant to your training goal.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#425
post #172

Earlier quoted context omitted.

also, China doing in IP what it's better at and way more experienced than USA - stealing.

This kind of blithe commentary is 20 years out of date and reminiscent of 1970s criticisms of the Japanese car industry.

gentle stroll through the aliexpress alleyway tells otherwise.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#426
post #374

Everyone is responding to the intellectual property issue, but isn't that the less interesting point? If Deepseek trained off OpenAI, then it wasn't trained from scratch for "pennies on the dollar" and isn't the Sputnik-like technical breakthrough that we've been hearing so much about. That's the news here. Or rather, the potential news, since we don't know if it's true yet.

[deleted]

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#427
I don't understand how OpenAI claims it would have happened. The weights are closed and as far as I read they are not complaining Deepseek hacked them and obtained the weight. So all they could do was to query OpenAI and generate test data. But how much did they query really - I would suppose it would require a huge amount done via an external, paid-for API? Is there any proof of this besides OpenAI saying it? Even if we suppose it is true, I suppose this must have happened via the API so they paid per token etc. So they paid for each and every token of training data. As I understand, the requester owns the copyright on what is generated by OpenAI's models and is free to do what they want.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#428
post #326
post #28

This smells very suspiciously like: someone who doesn't know anything about AI (possibly Sacks) demanding answers on R1 from someone who doesn't have any good ones (possibly Altman). "Uh, (sweating), umm, (shaking), they stole it from us! Yeah, look at this suspicious activity, that's why they had it so easy, we did all the hard work first!"

I think it's funny that OpenAI wants us to pay them to use their product to generate content but then sets the terms that they control how we use the content in generates for us. It takes someone like Deepseek to challenge that on our behalf or they will control most of the economy.

It’s quite ironic of them to claim that the only thing you cannot train on is another LLM output.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#429

This reminds me of the railroads, where once railroads were invented, there was a huge investment boom of eveyrone trying to make money of the railroads, but the competition brought the costs down where the railroads weren’t the people who generally made the money and got the benefit, but the consumers and regular businesses did and competition caused many to fail. AI is probably similar where the Moore’s law and adv…

> where the Moore’s law and advancement will eventually allow people to run open models locally

Probably won't be Moore's law (which is kind of slowing down) so much as architectural improvements (both on the compute side and the model side - you could say that R1 represents an architectural improvement of efficiency on the model side).

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#430

All the top level comments are basking in the irony of it, which is fair enough. But I think this changes the Deepseek narrative a bit. If they just benefited from repurposing OpenAI data, that's different than having achieved an engineering breakthrough, which may suggest OpenAI's results were hard earned after all.

Yeah what happens when we remove all financial incentive to fund groundbreaking science? It’s the same problem with pharmaceuticals and generics. It’s great when the price of drugs is low, but without perverse financial incentives no company is going to burn billions of dollars in a risky search for new medicines.

In this case, these cures (llms) are medicines in search for a disease to cure. I got Ai shoved everywhere, where I just want it to aid in my coding. Literally, that's it. They're also good at summarizing emails and similar things, but I know nobody who does that. I wouldn't trust an Ai reading and possibly hallucinate emails
Post reply on HN