Live data from Hacker News

OpenAI says it has evidence DeepSeek used its model to train competitor

ft.com

771–780 of 1001 posts

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#771
post #711

OpenAI is taking the position similar to that if you sell a cook book, people are not allowed to teach the recipes to their kids, or make better versions of them. That is absurd. Copyright law is designed to strike a balance between two issues. One the one hand, the creator’s personality that’s baked into the specific form of expression. And on the other hand, society’s interest in ideas being circulated, improved an…

Just to play devil's advocate, OAI can argue that they spent great effort creating and procuring annotated data. Such datasets are indeed their secret, and now DS gets them for free by distilling OAI's output. Besides, OAI's EULA explicitly forbids users from using the output of their API for model training. I'm not saying that OAI is right, of course. Just to present OAI's point of view.

Feist comprehensively rejected that argument under US copyright law, and the attempts in the late 90s to pass a law in response establishing a sui generis prohibition on copying databases also failed in the US. The EU did adopt a directive to that effect, which may be why there are no significant European search engines.

However, OpenAI and Google are far more politically influential than the lobbyists in the 90s, so it is likely to succeed.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#772
post #454
post #95

> “It’s also extremely hard to rally a big talented research team to charge a new hill in the fog together,” he added. “This is the key to driving progress forward.” Well I think DeepSeek releasing it open source and on an MIT license will rally the big talent. The open sourcing of a new technology has always driven progress in the past. The last paragraph too is where OpenAi seems to be focusing their efforts.. > we…

I'm willing to bet ''ban DeepSeek'' voices will start soon. Why compete, when you can just ban?

The fact it is out and improving day by day. Unsloth.ai is on a roll with their advancements. If DeepSeek is banned hundreds more will popup and change the data ever so slightly to skirt the ban. Pandora's box exploded on this one.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#773

OpenAI's models were trained on ebooks from a private ebook torrent tracker leeched en-mass during a free leech event by people who hated private torrent trackers and wanted to destroy their "economy." The books were all in epub format, converted, cleaned to plain text, and hosted on a public data hoarder site.

Have you got some support for this claim?

There's a lot of wild claims about, so while this is plausible it would be great if there were some evidence backing it.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#774

[flagged]

Claude will also often refer to itself as ChatGPT. This is not a 'China' problem.

Jesus Christ people!

If you stole my wallet because someone else stole your wallet it DOSE NOT MAKE IT OK.

Stealing is stealing. Stop excusing it.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#775

Earlier quoted context omitted.

If Deepseek trained off OpenAI, then it wasn't trained from scratch for "pennies on the dollar" If OpenAI trained on the intellectual property of others, maybe it wasn't the creativity breakthrough people claim? Oppositely If you say ChatGPT was trained on "whatever data was available", and you say Deepseek was trained "whatever data was available", then they sound pretty equivalent. All the rough consensus language…

I'm not an OpenAI apologist and don't like what they've done with other people's intellectual property but I think that's kind of a false equivalency. OpenAI's GPT 3.5/4 was a big leap forward in the technology in terms of functionality. DeepSeek-r1 isn't really a huge step forward in output, it's mostly comparable to existing models, one thing that is really cool about it is it being able to be trained from scratch…

How is it a “lie” for DeepSeek to train their data from ChatGPT but not if they train their data from all of Twitter and Reddit? Either way the training is 100x cheaper.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#776
post #773

OpenAI's models were trained on ebooks from a private ebook torrent tracker leeched en-mass during a free leech event by people who hated private torrent trackers and wanted to destroy their "economy." The books were all in epub format, converted, cleaned to plain text, and hosted on a public data hoarder site.

Have you got some support for this claim? There's a lot of wild claims about, so while this is plausible it would be great if there were some evidence backing it.

NYT claims that OpenAI trained on their material. They argue for copyright violation, although I think another argument might be breach of TOS in scraping the material from their website or archive.

The complaint filing has some references to some of the other training material used by OpenAI, but I didn't dig deeply in to what all of it was:

https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20...

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#777
post #320

I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…

Its a decent point if their models were not trained in isolation, but used o1 to improve it. But its rich from OpenAI to come complain DeepSeek or anyone else used their data for training. Get out fellow theives.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#779
post #504

Earlier quoted context omitted.

Why would it cast any doubt? If you can use o1 output to build a better R1. Then use R1 output to build a better X1... then a better X2.. XN, that just shows a method to create better systems for a fraction of the cost from where we stand. If it was that obvious OpenAI should have themselves done. But the disruptors did it. It hindsight it might sound obvious, but that is true for all innovations. It is all good stuf…

I think it would cast doubt on the narrative "you could have trained o1 with much less compute, and r1 is proof of that", if it turned out that in order to train r1 in the first place, you had to have access to bunch of outputs from o1. In other words, you had to do the really expensive o1 training in the first place. (with the caveat that all we have right now are accusations that DeepSeek made use of OpenAI data -…

> you had to do the really expensive o1 training in the first place

It is no better for OpenAI in this scenario either, any competitor can easily copy their expensive training without spending the same, i.e. there is a second mover advantage and no economic incentive to be the first one.

To put it another way, the $500 Billion Stargate investment will be worth just $5Billion once the models become available for consumption, because it only will take that much to replicate the same outcomes with new techniques even if the cold start needed o1 output for RL.

Post reply on HN