Hard to really have any sympathy for OpenAI's position when they're actively stealing content, ignoring requests to stop then spending huge amounts to get around sites running ai poisoning scripts, making it clear they'll still take your content regardless of if you consent to it.
Can someone with more expertise help me understand what I'm looking at here? https://crt.sh/?id=10106356492 It looks like Deepseek had a subdomain called "openai-us1.deepseek.com". What is a legitimate use-case for hosting an openai proxy(?) on your subdomain like this? Not implying anything's off here, but it's interesting to me that this OpenAI entity is one of the few subdomains they have on their site
OpenAI says it has evidence DeepSeek used its model to train competitor
891–900 of 1001 posts
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#892I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…
> This is obviously extremely silly, because that's exactly how OpenAI got all of its training data IANAL, but It is worth noting here that DeepSeek has explicitly consented to a license that doesn't allow them to do this . That is a condition of using the Chat GPT and the OpenAI API. Even if the courts affirm that there's a fair use defence for AI training, DeepSeek may still be in the wrong here, not because of cop…
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#893The cat is out of the bag. This is the landscape now, r1 was made in a post-o1 world. Now other models can distill r1 and so on. I don’t buy the argument that distilling from o1 undermines deep seek’s claims around expense at all. Just as open AI used the tools ‘available to them’ to train their models (eg everyone else’ data), r1 is using today’s tools. Does open AI really have a moral or ethical high ground here?
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#894Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#895Earlier quoted context omitted.
I understand they just used the API to talk to the OpenAI models. That... seems pretty innocent? Probably they even paid for it? OpenAI is selling API access, someone decided to buy it. Good for OpenAI! I understand ToS violations can lead to a ban. OpenAI is free to ban DeepSeek from using their APIs.
Sure, but I'm not interested in innocence. They can be as innocent or guilty as they want. But it means they didn't, via engineering wherewithal, reproduce the OpenAI capabilities from scratch. And originally that was supposed to be one of the stunning and impressive (if true) implications of the whole Deepseek news cycle.
who cares. even if the claim is true, does that make the open source model less attractive?
in fact, it implies that there is no moat in this game. openai can no longer maintain its stupid valuation, as other companies can just scrape its output and build better models at much lower costs.
everything points to the exact same end result - DeepSeek democratized AI, OpenAI's old business model is dead.
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#896I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…
> "DeepSeek trained on our outputs" I'm wondering how Deepseek could have made 100s of millions of training queries to OpenAI and not one person at OpenAI caught on.
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#897> “It’s also extremely hard to rally a big talented research team to charge a new hill in the fog together,” he added. “This is the key to driving progress forward.” Well I think DeepSeek releasing it open source and on an MIT license will rally the big talent. The open sourcing of a new technology has always driven progress in the past. The last paragraph too is where OpenAi seems to be focusing their efforts.. > we…
Actually the "our IP" argument is ridiculous. What they are doing is stealing data from all over the web, without people's consent for that data to be used in training ML models. If anything, then "Open"AI should be sued and forced to publish their whole product. The people should demand knowing exactly what is going on with their data. Also still an unresolved issue is how they will ever comply with a deletion reque…
But DeepSeek didn't use that presumably (since it's secret). They definitely can't argue that using copyrighted material for training is fine, but using output from other commercial models isn't. That's too inconsistent.
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#898Earlier quoted context omitted.
Why would it cast any doubt? If you can use o1 output to build a better R1. Then use R1 output to build a better X1... then a better X2.. XN, that just shows a method to create better systems for a fraction of the cost from where we stand. If it was that obvious OpenAI should have themselves done. But the disruptors did it. It hindsight it might sound obvious, but that is true for all innovations. It is all good stuf…
They're standing on the shoulders of giants, not only in terms of re-using expensive computing power almost for free by using the outputs of expensive models. It's a bit of a tradition in that country, also in manufacturing.
Re: OpenAI says it has evidence DeepSeek used its model to train competitor
#899I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating…
This is a fascinating development because AI models may turn out to be like pharmaceuticals. The first pill costs $500 million to make, the second one costs pennies.
Besides deals with insurance companies and governments, one of the ways that they are still able to pull this is convincing everyone that it's too dangerous to play with this at home or buying it from an Asian supplier.
At least with software we had until now a way to build and run most things without requiring dedicated super expensive equipment. OpenAI pulled a big Pharma move but hopefully there will be enough disruptors to not let them continue it.