Live data from Hacker News

Moonshot serves Claude instead of Kimi and collects exchanges for model training

twitter.com

11–20 of 74 posts

Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training

#11

There is absolutely nothing ethically wrong with other models using GPT and Claude for training. Heck, GPT-Sol even helped train GPT-Luna. There also is nothing wrong with using customer data for training, since millions if not billions of other users benefit from it. Again, even the big players recognize its relevance and do it.

[flagged]

Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training

#12
post #11

There is absolutely nothing ethically wrong with other models using GPT and Claude for training. Heck, GPT-Sol even helped train GPT-Luna. There also is nothing wrong with using customer data for training, since millions if not billions of other users benefit from it. Again, even the big players recognize its relevance and do it.

[flagged]

Violating terms of use is not illegal.

Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training

#13
post #5

Claude/OpenAI etc have taken and continue to take literally all data from the internet, printed books, image, audio, and video humans have created in all of existence without permission to train their models. But when same AI company gets "distilled" or it's own AI-generated content used to train other models, it's suddenly immoral or illegal?

People keep saying this without attribution. Meta and others got caught, and brought into court, over using torrented files. But those models trained on that data have long since been retired and replaced with new models based on new from-scratch training runs. OpenAI, Anthropic, etc all pay studios, newspapers, Reddit and others for access to data for training. They scrape the open web, but if that's illegal a court hasn't said so. The open web is open, after all. And they don't seem to be stealing books, they seem to be buying physical copies and scanning. Seems legit, that's what a human would do to learn from a book. They also pay big bucks for commercially curated data and training sets.

Just feels like there's enormous CCP effort to put their labs on equal moral footing with everyone else when it's not demonstrably the case. They want the West to hate themselves so we're happy to squander our technological lead.

Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training

#15
post #13
post #5

Claude/OpenAI etc have taken and continue to take literally all data from the internet, printed books, image, audio, and video humans have created in all of existence without permission to train their models. But when same AI company gets "distilled" or it's own AI-generated content used to train other models, it's suddenly immoral or illegal?

People keep saying this without attribution. Meta and others got caught, and brought into court, over using torrented files. But those models trained on that data have long since been retired and replaced with new models based on new from-scratch training runs. OpenAI, Anthropic, etc all pay studios, newspapers, Reddit and others for access to data for training. They scrape the open web, but if that's illegal a court…

> they seem to be buying physical copies and scanning. Seems legit, that's what a human would do to learn from a book

There’s still the open question on learn vs copy/mimic/repeat.

As a human, I can read a book I bought. I’m definitely not allowed to scan it and post its pages online and upload them to an archive of scanned PDFs without the authors’ and publishers’ permission.

IIRC the Meta legal case wasn’t even about LLMs, they just torrented and shared pirated files, whether with strangers or among employees. Those may or may not have been later used for training, but it was already illegal to just share among employees.

Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training

#16
The open question here is how is Anthropic retaining these exchanges?

If Moonshot is using the API, normally Anthropic would not retain the exchanges, at least that's the promise. If Anthropic is consistently retaining all exchanges from a class of customers because they're "flagged" but not notifying those customers, how can an average customer trust it won't happen to them?

If Moonshot is buying accounts and using those rather than the API, wouldn't they set the "no training on my data" flag in the settings so as to go undetected for longer? If so, we get back to the question of why would Anthropic be retaining the exchanges?

Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training

#17
post #7

There is absolutely nothing ethically wrong with other models using GPT and Claude for training. Heck, GPT-Sol even helped train GPT-Luna. There also is nothing wrong with using customer data for training, since millions if not billions of other users benefit from it. Again, even the big players recognize its relevance and do it.

They are using stolen api keys and credentials.

They probably have no choice because they're not allowed to pay for its use. If they were allowed to pay for a Claude or Codex subscription, they probably would.

Also, it is essential for the model's responses to be without censorship. In reality, responses sent to China by flagship models are known to be weaker.

Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training

#18
post #13
post #5

Claude/OpenAI etc have taken and continue to take literally all data from the internet, printed books, image, audio, and video humans have created in all of existence without permission to train their models. But when same AI company gets "distilled" or it's own AI-generated content used to train other models, it's suddenly immoral or illegal?

People keep saying this without attribution. Meta and others got caught, and brought into court, over using torrented files. But those models trained on that data have long since been retired and replaced with new models based on new from-scratch training runs. OpenAI, Anthropic, etc all pay studios, newspapers, Reddit and others for access to data for training. They scrape the open web, but if that's illegal a court…

While many things might be legal (or hasn't been ruled clearly illegal yet) /under different jurisdictions, we are still in the process of figuring out what we accept as ethical. As the output of models can't be easily copyrighted, destillation is equally disputed. Particularly if the primary model interaction was not destillation (as in this case) IMHO it will be legally quite difficult to restrict secondary use for training. In the end we have to find a legislation and probably even international treaties that account for the fact that classical copyright is beginning to become an obsolete concept.

Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training

#19
post #13
post #5

Claude/OpenAI etc have taken and continue to take literally all data from the internet, printed books, image, audio, and video humans have created in all of existence without permission to train their models. But when same AI company gets "distilled" or it's own AI-generated content used to train other models, it's suddenly immoral or illegal?

People keep saying this without attribution. Meta and others got caught, and brought into court, over using torrented files. But those models trained on that data have long since been retired and replaced with new models based on new from-scratch training runs. OpenAI, Anthropic, etc all pay studios, newspapers, Reddit and others for access to data for training. They scrape the open web, but if that's illegal a court…

Just because the material was legally required doesn't mean you can do anything you want with it. I can't (legally) buy a physical book, scan it, and put the scan on my web site. It seems to me that an LLM is a derived work of the training materials that went into it, and thus needs permission from the copyright holders.

But OK, the law seems to disagree with me there. But OK, let's say it's fine for AI companies to train their models on copyrighted content as long as they didn't torrent it or whatever. What then makes it illegal, or morally wrong, to do the same thing with their competitors' model outputs? Why is it OK for Anthropic to scrape this comment and feed it into their system, but not OK for Moonshot to scrape the output of Anthropic's system and feed it into theirs?

Post reply on HN