There is absolutely nothing ethically wrong with other models using GPT and Claude for training. Heck, GPT-Sol even helped train GPT-Luna. There also is nothing wrong with using customer data for training, since millions if not billions of other users benefit from it. Again, even the big players recognize its relevance and do it.
Moonshot serves Claude instead of Kimi and collects exchanges for model training
11–20 of 74 posts
Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training
#12There is absolutely nothing ethically wrong with other models using GPT and Claude for training. Heck, GPT-Sol even helped train GPT-Luna. There also is nothing wrong with using customer data for training, since millions if not billions of other users benefit from it. Again, even the big players recognize its relevance and do it.
[flagged]
Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training
#13Claude/OpenAI etc have taken and continue to take literally all data from the internet, printed books, image, audio, and video humans have created in all of existence without permission to train their models. But when same AI company gets "distilled" or it's own AI-generated content used to train other models, it's suddenly immoral or illegal?
Just feels like there's enormous CCP effort to put their labs on equal moral footing with everyone else when it's not demonstrably the case. They want the West to hate themselves so we're happy to squander our technological lead.
Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training
#14The only problem with it is they might have worse security than Anthropic and your personal info gets leaked, but I don't think it can happen that easy.
Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training
#15Claude/OpenAI etc have taken and continue to take literally all data from the internet, printed books, image, audio, and video humans have created in all of existence without permission to train their models. But when same AI company gets "distilled" or it's own AI-generated content used to train other models, it's suddenly immoral or illegal?
People keep saying this without attribution. Meta and others got caught, and brought into court, over using torrented files. But those models trained on that data have long since been retired and replaced with new models based on new from-scratch training runs. OpenAI, Anthropic, etc all pay studios, newspapers, Reddit and others for access to data for training. They scrape the open web, but if that's illegal a court…
There’s still the open question on learn vs copy/mimic/repeat.
As a human, I can read a book I bought. I’m definitely not allowed to scan it and post its pages online and upload them to an archive of scanned PDFs without the authors’ and publishers’ permission.
IIRC the Meta legal case wasn’t even about LLMs, they just torrented and shared pirated files, whether with strangers or among employees. Those may or may not have been later used for training, but it was already illegal to just share among employees.
Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training
#16If Moonshot is using the API, normally Anthropic would not retain the exchanges, at least that's the promise. If Anthropic is consistently retaining all exchanges from a class of customers because they're "flagged" but not notifying those customers, how can an average customer trust it won't happen to them?
If Moonshot is buying accounts and using those rather than the API, wouldn't they set the "no training on my data" flag in the settings so as to go undetected for longer? If so, we get back to the question of why would Anthropic be retaining the exchanges?
Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training
#17There is absolutely nothing ethically wrong with other models using GPT and Claude for training. Heck, GPT-Sol even helped train GPT-Luna. There also is nothing wrong with using customer data for training, since millions if not billions of other users benefit from it. Again, even the big players recognize its relevance and do it.
They are using stolen api keys and credentials.
Also, it is essential for the model's responses to be without censorship. In reality, responses sent to China by flagship models are known to be weaker.
Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training
#18Claude/OpenAI etc have taken and continue to take literally all data from the internet, printed books, image, audio, and video humans have created in all of existence without permission to train their models. But when same AI company gets "distilled" or it's own AI-generated content used to train other models, it's suddenly immoral or illegal?
People keep saying this without attribution. Meta and others got caught, and brought into court, over using torrented files. But those models trained on that data have long since been retired and replaced with new models based on new from-scratch training runs. OpenAI, Anthropic, etc all pay studios, newspapers, Reddit and others for access to data for training. They scrape the open web, but if that's illegal a court…
Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training
#19Claude/OpenAI etc have taken and continue to take literally all data from the internet, printed books, image, audio, and video humans have created in all of existence without permission to train their models. But when same AI company gets "distilled" or it's own AI-generated content used to train other models, it's suddenly immoral or illegal?
People keep saying this without attribution. Meta and others got caught, and brought into court, over using torrented files. But those models trained on that data have long since been retired and replaced with new models based on new from-scratch training runs. OpenAI, Anthropic, etc all pay studios, newspapers, Reddit and others for access to data for training. They scrape the open web, but if that's illegal a court…
But OK, the law seems to disagree with me there. But OK, let's say it's fine for AI companies to train their models on copyrighted content as long as they didn't torrent it or whatever. What then makes it illegal, or morally wrong, to do the same thing with their competitors' model outputs? Why is it OK for Anthropic to scrape this comment and feed it into their system, but not OK for Moonshot to scrape the output of Anthropic's system and feed it into theirs?