Live data from Hacker News

Anthropic says Alibaba illicitly extracted Claude AI model capabilities

reuters.com

821–830 of 1001 posts

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#821
post #785

Earlier quoted context omitted.

> Distillation is NOT an attack. From the article - > 28.8 million exchanges with Claude through almost 25,000 fraudulent accounts wouldn't that be considered an attack? Not sure what I'm missing here.

It's merely a ToS violation.

Exactly, calling it “illicit” is funny. Your ToS isn’t law.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#822
post #785

Earlier quoted context omitted.

> These complaints of distillation are inflating the problem to make it sound worse than it is Unfortunately, the Reuters piece itself is complicit in this dramatization. The lede paragraph parrots Anthropic's talking point that distillation is an "attack", without using quotes that would alert the reader that this framing is a corporate talking point. Distillation is NOT an attack.

> Distillation is NOT an attack. From the article - > 28.8 million exchanges with Claude through almost 25,000 fraudulent accounts wouldn't that be considered an attack? Not sure what I'm missing here.

Just sending a request to a service does not constitute an "attack". It seems that what Anthropic mean by "fraudulent account" is probably just one violating their terms of service - misuse of a subscription account, and/or the presumed nature of what the user was trying to do.

I guess Anthropoic would regard any developer using their subscription plan with OpenCode to be operating a "fraudulent account", maybe an "attacker" too. Now we know how they think of anyone using Claude to develop software competing with Anthropic. Only an "attacker" would want to vibe code their own harness, or god forbid want to learn how to build/train an LLM.

Of course Anthropic's wording is intended to be deliberately provocative, since they are trying to manipulate the US government into shutting down the Chinese competition.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#823

“Distillation attack” are we joking here. If anything these models should be compelled to be public since they have been trained off public data. What an absurd overreach to call this an attack. It’s clear they are scapegoating national security and China at this point to build an anti-competitive moat. I generally really like Anthropic’s work and models but stuff like this scares me for the future. We are positionin…

> If anything these models should be compelled to be public since they have been trained off public data I'm starting to come around to this idea TBH. For a while my position was: "these companies have invested billions into training these models, therefore they should be able to control them and profit off them" but looking deeper at where they got their training data, my view is starting to shift. IMHO I feel like…

labs invest multiple billion dollars a year each in private data, and that number is growing. internet training data is not where frontier capabilities come from, this view is outdated

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#824

“Distillation attack” are we joking here. If anything these models should be compelled to be public since they have been trained off public data. What an absurd overreach to call this an attack. It’s clear they are scapegoating national security and China at this point to build an anti-competitive moat. I generally really like Anthropic’s work and models but stuff like this scares me for the future. We are positionin…

> If anything these models should be compelled to be public since they have been trained off public data.

Isn't that a bit like saying if you read books in a public library to pick up a new skill you should work for free?

> What an absurd overreach to call this an attack.

Would it be an attack to take your meal by force if you used a public recipe to prepare the meal?

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#825
post #807

“Distillation attack” are we joking here. If anything these models should be compelled to be public since they have been trained off public data. What an absurd overreach to call this an attack. It’s clear they are scapegoating national security and China at this point to build an anti-competitive moat. I generally really like Anthropic’s work and models but stuff like this scares me for the future. We are positionin…

The core of the training data is public, but the part that actually makes these models smart came from (pretty highly-paid) experts via platforms like Mercor. Claude didn't magically learn to write good code by reading all of GitHub - humans trained it in that, more or less manually.

No, they do RLVR (reinforcement learning with verifiable rewards) like everyone else. And probably use claude data too, with human in the loop and tool feedback.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#826
post #819

“Distillation attack” are we joking here. If anything these models should be compelled to be public since they have been trained off public data. What an absurd overreach to call this an attack. It’s clear they are scapegoating national security and China at this point to build an anti-competitive moat. I generally really like Anthropic’s work and models but stuff like this scares me for the future. We are positionin…

> If anything these models should be compelled to be public since they have been trained off public data. What an absurd overreach to call this an attack. > It’s clear they are scapegoating national security and China at this point to build an anti-competitive moat. If all that is required to train these models is public data, why can't Alibaba just use that? The fact that Alibaba has to resort to scraping Claude sug…

This feels more nuanced than you are giving it credit for? Much of the training data that was available has been withdrawn, atleast for OpenAI we know that much of the training data was garnered in less-than above the board methods

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#827

Here's what is happening: Chinese resellers are offering Claude tokens at 70-90% below official Anthropic API prices. They achieve this by reselling capacity from pooled Claude Max accounts, payments fraud, and also reselling the model output & reasoning chains to various Chinese labs. They are subsidizing model access in exchange for user logs and reasoning traces, which they then sell as training data, allowing the…

I didn't connect the reseller pricing to DS and GLM prices until you explained it. Very good observation. Deepseek v4 pro in particular is priced so low that it's hard to imagine that they have any margin. 0.76/1.52 for a 1.6T param model leaves very little margin. Even the domestic providers on Openrouter are multiples of the price https://openrouter.ai/deepseek/deepseek-v4-pro

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#829

“Distillation attack” are we joking here. If anything these models should be compelled to be public since they have been trained off public data. What an absurd overreach to call this an attack. It’s clear they are scapegoating national security and China at this point to build an anti-competitive moat. I generally really like Anthropic’s work and models but stuff like this scares me for the future. We are positionin…

Should Google search index be forced to be public too?

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#830

Earlier quoted context omitted.

Reuters is probably the most rigorous news agency in the world. > it said was the largest known attack > Anthropic said in the letter it was supportive of the U.S. government's efforts to combat the attacks both times the word "attack" appears it's clearly stated that the word was used by the company, it's a direct company quote. actually putting it into quotes would be editorializing > Unfortunately, the Reuters pie…

> how would you feel if somebody quoting you would turn your word dramatization into "dramatization" because they don't agree with your assesment This is exactly what news agency should be doing though. When the dude showed up to Comet Pizza to look for Hillary Clinton or whatever, do you figure they should've printed "Local hero saves children from predatory cabal"?

I want them to report the facts, not their opinions.

Reporting that corporate called it attacks is good. I do prefer direct quotes.

However, when they quote one word, the journalists are inserting their own opinion about it. I want to make my own opinions based on the facts. I don't need the reporter to draw the conclusions for me.

Post reply on HN