Live data from Hacker News

OpenAI says it has evidence DeepSeek used its model to train competitor

ft.com

381–390 of 1001 posts

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#381

Earlier quoted context omitted.

It's about as illegal as the billions, if not trillions of IPs that ClosedAI infringed to train their own data without consent. Not that they're alone, and I personally don't mind that AI companies do it, but it's still amusing when they get this annoyed at others doing the same thing to them.

I think they had the advantage of being ahead of the law in this regard. To my knowledge, reading copywritten material isn't (or wasn't illegal) and remains a legal grey area. Distilling weights from prompts and responses is even more of a legal grey area. The legal system cannot respond quickly to such technological advancements so things necessarily remain a wild west until technology reaches the asymptotic portion…

The main issue is that they've had plenty of instances where the LLM outputted copyrighted content verbatim, like it happened with the New York Times and some book authors. And then there's DALL-E, which is baked into ChatGPT and before all the guardrails came up, was clearly trained on copyrighted content to the point it had people's watermarks, as well as their styles, just like Stable Diffusion mixes can do (if you don't prompt it out).

Like you've put, it's still a somewhat gray area, and I personally have nothing against them (or anyone else) using copyrighted content to train models.

I do find it annoying that they're so closed-off about their tech when it's built on the shoulders of openness and other people's hard work. And then they turn around and throw Issy fits when someone copies their homework, allegedly.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#382
I heard they were just “democratizing” llm and ai development.

Yesterday the industry crushed pianos and tools and bicycles and guitars and violins and paint supplies and replaced them with a tablet computer.

Tomorrow we can replace craven venture capitalists and overfed corporate bodies with incestuous LLM’s and call it all a day.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#385
post #374

Everyone is responding to the intellectual property issue, but isn't that the less interesting point? If Deepseek trained off OpenAI, then it wasn't trained from scratch for "pennies on the dollar" and isn't the Sputnik-like technical breakthrough that we've been hearing so much about. That's the news here. Or rather, the potential news, since we don't know if it's true yet.

But it does mean moat is even less defensible for companies whose fortunes are tied to their foundation models having some performance edge, and a shift in the kinds of hardware used for inference (smaller, closer to the edge.)

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#386
post #4

How would they prove they used it’s model. I would be curious to know their methodology. Also what legal actions OpenAI can take? can DeepSeek be banned in US?

They might show DeepSeek's model calling itself ChatGPT, which users have already alleged. Same as how Cisco proved Huawei was stealing router code.

Except in this case, nothing was stolen, unless they want to call ChatGPT's own training on source data theft too.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#389
post #374

Everyone is responding to the intellectual property issue, but isn't that the less interesting point? If Deepseek trained off OpenAI, then it wasn't trained from scratch for "pennies on the dollar" and isn't the Sputnik-like technical breakthrough that we've been hearing so much about. That's the news here. Or rather, the potential news, since we don't know if it's true yet.

That's not correct. First of all, training off of data generated by another AI is generally a bad idea because you'll end up with a strictly less accurate model (usually). But secondly, and more to your point, even if you were to use training data from another model, YOU STILL NEED TO DO ALL THE TRAINING.

Using data from another model won't save you any training time.

Re: OpenAI says it has evidence DeepSeek used its model to train competitor

#390
It seems to be undermined by the same principle that says that going into a library and reading a book there is not stealing when you walk out with the knowledge from the book.

OpenAI seems to feel that way about the their use of copyrighted material: since they didn't literally make a copy of the source material, it's totally fair game. It seems like this is the same argument that protects DeepSeek if indeed they did this. And why not, reading a lot of books from the library is a way to get smarter, and ostensibly the point of libraries

Post reply on HN