There's two basic kinds of distillation: 1) the massive [and dumb] method where you ask a question and use the answer as reinforcement (Black Box), and 2) more targeted distillation where you use one model to directly inform/train/guide another model (RLAIF). The latter is basically fine-tuning the model with direction from another model. Thousands of businesses do this every day to fine-tune. This is almost certainl…
Anthropic says Alibaba illicitly extracted Claude AI model capabilities
761–770 of 1001 posts
Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities
#762There's two basic kinds of distillation: 1) the massive [and dumb] method where you ask a question and use the answer as reinforcement (Black Box), and 2) more targeted distillation where you use one model to directly inform/train/guide another model (RLAIF). The latter is basically fine-tuning the model with direction from another model. Thousands of businesses do this every day to fine-tune. This is almost certainl…
> These complaints of distillation are inflating the problem to make it sound worse than it is Unfortunately, the Reuters piece itself is complicit in this dramatization. The lede paragraph parrots Anthropic's talking point that distillation is an "attack", without using quotes that would alert the reader that this framing is a corporate talking point. Distillation is NOT an attack.
Any reasonable company would be pissed if a competitor, especially at Ali Baba's size, leveraged that company's R&D to compete. It is in this sense, a corporate attack.
If you want to roll your eyes at distillation concerns, you might need to excuse Anthropic for originally using pirated material to train their models.
Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities
#763The hypocrisy of Anthropic complaining about "illicitly extracting its Claude AI model capabilities" and supporting the White House's accusation of China "stealing U.S. AI labs' intellectual property on an industrial scale" is hilarious. Anthropic, OpenAI, Google, Microsoft, et al trained their models by ignoring the rights of copyright holders when harvesting whatever content they could. Now one of them is crying fo…
Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities
#764If you've invested in expensive capabilities training, of course you don't want this, so it's in Anthropic's economic interest to hinder it however they can, and that's enough to explain their behaviour here.
Anthropic seems to genuinely care about safety though, which for the rest of us means not having models that enabling easier cyberattacks, targeted scams, and the rarer but more severe risks like people trying to create and release new pathogens. This means walking a tight line, especially as models become more capable, and often wrapping a model in layers of defences against misuse.
If those capabilities transfer to a closed competitor model, all bets are off in terms of whether the competitor will apply the same defences.
If those capabilities transfer to an open weight model, not only will there be no ring of defences around the model, any defences you put into the model itself can easily be stripped away. So although it's nice to have capable open models, it will increasingly bad for us all if open models keep fast-following closed model capabilities as they have been, at least until we have solved the active research problem of keeping them safe.
This is all to say that, however you might feel about Anthropic, we might still prefer that they can deter this kind of distillation for now.
Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities
#765This kind of systematic distillation by a competitor can allow them to fast-follow you and pick up capabilities. If you've invested in expensive capabilities training, of course you don't want this, so it's in Anthropic's economic interest to hinder it however they can, and that's enough to explain their behaviour here. Anthropic seems to genuinely care about safety though, which for the rest of us means not having m…
There are sometimes false positives but when I give Kimi’s report to the frontier models they more often than not confirm they are valid security issues but didn’t find them themselves.
Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities
#766Earlier quoted context omitted.
> These complaints of distillation are inflating the problem to make it sound worse than it is Unfortunately, the Reuters piece itself is complicit in this dramatization. The lede paragraph parrots Anthropic's talking point that distillation is an "attack", without using quotes that would alert the reader that this framing is a corporate talking point. Distillation is NOT an attack.
Distillation may not be an attack, but it is a ToS violation and could be seen as IP theft. Any reasonable company would be pissed if a competitor, especially at Ali Baba's size, leveraged that company's R&D to compete. It is in this sense, a corporate attack. If you want to roll your eyes at distillation concerns, you might need to excuse Anthropic for originally using pirated material to train their models.
Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities
#767Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities
#768If not, then we should look at Alibaba, but we should look at Anthropic as well.
Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities
#769Earlier quoted context omitted.
Wait, so is your theory mutually exclusive to Anthropic's claims of "theft of capabilities"?
No, this reseller 中转站 thing is basically a loss leader for certain chinese ai labs to distill claude with verified human input.
Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities
#770Earlier quoted context omitted.
The compute deficit of Chinese Ai companies is real, and it IS THE ONLY competitive advantage that Western companies have. The only way the U.S. keeps that edge is to prevent distillation. The only way Chinese companies can make up for the deficit in compute is to distill. There innovation in great supply on every side of the Ocean. Its about the chips. And in terms of national security, for the U.S., and for China,…
Define compute deficit? They've been bringing out open weight models competitive with frontier models. How could they do that if they had a compute deficit?
I'm using GLM-5.2 daily for my own stuff, and during Chinese business hours, specially on their afternoon, it's a festival of rate limits.