Claude/OpenAI etc have taken and continue to take literally all data from the internet, printed books, image, audio, and video humans have created in all of existence without permission to train their models. But when same AI company gets "distilled" or it's own AI-generated content used to train other models, it's suddenly immoral or illegal?
Moonshot serves Claude instead of Kimi and collects exchanges for model training
21–30 of 74 posts
Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training
#22Claude/OpenAI etc have taken and continue to take literally all data from the internet, printed books, image, audio, and video humans have created in all of existence without permission to train their models. But when same AI company gets "distilled" or it's own AI-generated content used to train other models, it's suddenly immoral or illegal?
People keep saying this without attribution. Meta and others got caught, and brought into court, over using torrented files. But those models trained on that data have long since been retired and replaced with new models based on new from-scratch training runs. OpenAI, Anthropic, etc all pay studios, newspapers, Reddit and others for access to data for training. They scrape the open web, but if that's illegal a court…
Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training
#23Earlier quoted context omitted.
People keep saying this without attribution. Meta and others got caught, and brought into court, over using torrented files. But those models trained on that data have long since been retired and replaced with new models based on new from-scratch training runs. OpenAI, Anthropic, etc all pay studios, newspapers, Reddit and others for access to data for training. They scrape the open web, but if that's illegal a court…
Just because the material was legally required doesn't mean you can do anything you want with it. I can't (legally) buy a physical book, scan it, and put the scan on my web site. It seems to me that an LLM is a derived work of the training materials that went into it, and thus needs permission from the copyright holders. But OK, the law seems to disagree with me there. But OK, let's say it's fine for AI companies to…
distilling another llm is a clear tos violation but no one really knows how much teeth those have. financially probably none all they can do is whack a mole on the accounts doing it which won’t work.
so they’re trying to lobby copyright changes i guess; unlikely to succeed as doing so would also make all search engines illegal
Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training
#24The open question here is how is Anthropic retaining these exchanges? If Moonshot is using the API, normally Anthropic would not retain the exchanges, at least that's the promise. If Anthropic is consistently retaining all exchanges from a class of customers because they're "flagged" but not notifying those customers, how can an average customer trust it won't happen to them? If Moonshot is buying accounts and using…
I would like to briefly draw your attention to this part of the Terms of Service [0]:
> We may use Materials to provide, maintain, and improve the Services and to develop other products and services, including training our models, unless you opt out of training through your account settings. Even if you opt out, we will use Materials for model training when: (1) you provide Feedback to us regarding any Materials, or (2) your Materials are flagged for safety review to improve our ability to detect harmful content, enforce our policies, or advance our safety research.
From the privacy policy [1]:
> We use your personal data for the following purposes:
> To prevent and investigate fraud, abuse, and violations of our Usage Policy, unlawful or criminal activity, unauthorized access to or use of personal data or Anthropic systems and networks, to protect our rights and the rights of others, to protect your safety or that of any other person, and to meet legal, governmental and institutional policy obligations;
Anthropic does offer a truly no data retention service, which is their "zero-data retention" agreement, which you have to sign and provide more documentation for than simply using either a normal consumer or API account. I doubt Moonshot did that, so by their terms, they can retain your data and use it for training if it's part of abuse prevention.
Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training
#25Earlier quoted context omitted.
Just because the material was legally required doesn't mean you can do anything you want with it. I can't (legally) buy a physical book, scan it, and put the scan on my web site. It seems to me that an LLM is a derived work of the training materials that went into it, and thus needs permission from the copyright holders. But OK, the law seems to disagree with me there. But OK, let's say it's fine for AI companies to…
you can buy a book, scan it, and upload the counts of every letter, distribution of apostophies, use it as the input to some convoluted process to produce weights or a search index though. They got slapped for illegally obtaining the files, not for producing derivative works of them. distilling another llm is a clear tos violation but no one really knows how much teeth those have. financially probably none all they c…
Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training
#26Claude/OpenAI etc have taken and continue to take literally all data from the internet, printed books, image, audio, and video humans have created in all of existence without permission to train their models. But when same AI company gets "distilled" or it's own AI-generated content used to train other models, it's suddenly immoral or illegal?
It is moreover established that training weights on basically anything is legitimate use.
Why repeat lie after lie like this? I don’t like LLM mania either but after reading the ten millionth mind-numbing insult to HN intelligence like this I have to think my mother gave better instruction.
Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training
#27Framing learning from observation as somehow bad. Something literally everybody is doing, model and human alike. It's the process this whole industry is built on. Pretending that this is bad because the other people are also doing what you have been doing, is hypocritical and childish.
Meanwhile I'm sitting here looking at some "Flibbertigittering" spinner like some sort of caveman, unable to steer the model when it misinterprets my ambigious prompt because I'm not allowed to see 80% of the output tokens I'm paying for.
Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training
#28Claude/OpenAI etc have taken and continue to take literally all data from the internet, printed books, image, audio, and video humans have created in all of existence without permission to train their models. But when same AI company gets "distilled" or it's own AI-generated content used to train other models, it's suddenly immoral or illegal?
People keep saying this without attribution. Meta and others got caught, and brought into court, over using torrented files. But those models trained on that data have long since been retired and replaced with new models based on new from-scratch training runs. OpenAI, Anthropic, etc all pay studios, newspapers, Reddit and others for access to data for training. They scrape the open web, but if that's illegal a court…
Is it theft? Well, no. There's no authentication bypass here, no Claude model leak. At best it is violating the terms of use, kind of like how it is violating the terms of use to scrape many websites that AI scrapers scraped.
Is it immoral? Why would it be, exactly? Distillation is not a forbidden technique with moral implications. In fact, there is quite compelling evidence that Anthropic themselves were distilling from OpenAI in early Claude models. It helped them bootstrap if nothing else. There is no special moral code that makes distillation forbidden any more than training off of people's works without permission, or even express non-consent, is forbidden.
Really the more concerning aspect of this is the deception of using Kimi and expecting Kimi output and getting Claude instead, but I would like some independent confirmation that this is even something Moonshot really did before raking them over the coals, rather than just assuming it's true because Anthropic said so. How exactly did they figure out, considering ZDR? It deserves more information.
I do agree that there is a tendency for people to justify CCP human rights violations by trying to equate them to much lesser but similarly shaped transgressions from Western governments, but that's an unrelated issue entirely. The story regarding distillation is consistent: Sorry, but I can't afford enough tiny violins to express my lack of giving a shit. I harbor no ill will, I truly hope the golden parachutes that Sam and Dario fly out on are adorned with the finest materials.
Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training
#29Claude/OpenAI etc have taken and continue to take literally all data from the internet, printed books, image, audio, and video humans have created in all of existence without permission to train their models. But when same AI company gets "distilled" or it's own AI-generated content used to train other models, it's suddenly immoral or illegal?
Pretending to customers like they are serving Kimi while actually proxying Claude is however a bad thing to do, bordering on fraud. I can see at least three issues
- data privacy. I might not want Anthropic to have my data. Even agreeing to Moonshot training on the data is not the same as giving them the right to send it to a completely different jurisdiction to do whatever
- it distorts model performance. If I evaluate Kimi based on their API, but the benchmarks happen to get sent to Claude, that gives the wrong impression. Or the other way around, if I evaluate based on the real Kimi and then my "production" requests get sent to Claude
- people build applications about the behavior of the model they are targeting. The models are nondeterministic, but they do have "flavors" and typical response patterns. Just switching out the models for a completely different model family is likely to lead to unexplained breakage
All of those apply to hobby applications and personal use just as much as to professional use
Re: Moonshot serves Claude instead of Kimi and collects exchanges for model training
#30The open question here is how is Anthropic retaining these exchanges? If Moonshot is using the API, normally Anthropic would not retain the exchanges, at least that's the promise. If Anthropic is consistently retaining all exchanges from a class of customers because they're "flagged" but not notifying those customers, how can an average customer trust it won't happen to them? If Moonshot is buying accounts and using…
> If Moonshot is buying accounts and using those rather than the API, wouldn't they set the "no training on my data" flag in the settings so as to go undetected for longer? If so, we get back to the question of why would Anthropic be retaining the exchanges? I would like to briefly draw your attention to this part of the Terms of Service [0]: > We may use Materials to provide, maintain, and improve the Services and t…
If the criteria for retention justifiable as safety at Anthropic are cast very broadly in actual practice, then it opens cans of worms for a lot of customers in complying with privacy regimes, customer IP protection, etc. It also raises the question of whether at least some subset of customers from certain geos (e.g., China likely, but perhaps other geos) may just automatically have everything flagged for retention, despite the terms, based on collective abuse risk flagging of their geo.
While ZDR is available, it's not available to all types of customers and is priced at a higher tier. Undoubtedly though this may be a wakeup call to some customers who didn't realize they may need ZDR to get on the ZDR bandwagon.