Live data from Hacker News

Anthropic says Alibaba illicitly extracted Claude AI model capabilities

reuters.com

481–490 of 1001 posts

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#481
post #477

I don't understand. If they are simply using our API and paying for tokens, it's called a "transaction" and not "attack". The user is our customer who is supporting our business by buying our services. And we call them attackers. We happily make money by selling our services, and then call it as attack. Back in the day, an "attack" was supposed to mean be someone acquiring our assets without paying for them or withou…

Did Anthropic 'attack' all those authors it was forced to pay $1.5bn to for using their work without permission?

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#482
post #340
post #246

Earlier quoted context omitted.

The US labs do seem to have announced a lot of licensing deals though, and are buying things today due to the previous lawsuits. At what point will we be better to support a lab that pays (some) licenses today vs the ones that pay none? Some of the deals are in the hundreds of millions, so I suspect licensing is over a billion today? (Pure guess). That might become a big disadvantage in a price (or content) war.

At the very least the public should receive full open-weight open-source models in return for their transgressions. Failing that, may I suggest the guillotine?

[deleted]

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#483
post #337
post #334

> The strike by Alibaba is described as a "distillation" effort, which Anthropic has said involves training a less capable model on the outputs of a stronger one. I don't see what's wrong about this. > Anthropic said the campaign was conducted between April 22 and June 5, 2026, and generated more than 28.8 million exchanges with Claude through almost 25,000 fraudulent accounts. What makes the accounts fraudulent? If…

Because Anthropic has terms of service with more stipulations than just "you must pay and can use the service for any purpose"?

So does a lot of the owners of data that Anthropic used for training. Anthropic preceeded to ignore said terms under the guise of fair use. Yet now they cry faul? Cry me a river.

To be clear: In principle I'm on Anthropic's side here. But Anthropic et al. have been very clear that they want to take a huge dump on those principles, so here we are.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#484
Sounds like just a case of pirates "illicitly" stealing from pirates. I don't really see anything ethically questionable there. I wonder if US corps will ever come out about all the resources used to train the original models and who they actually asked for permission when collecting data.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#485

Earlier quoted context omitted.

> These complaints of distillation are inflating the problem to make it sound worse than it is Unfortunately, the Reuters piece itself is complicit in this dramatization. The lede paragraph parrots Anthropic's talking point that distillation is an "attack", without using quotes that would alert the reader that this framing is a corporate talking point. Distillation is NOT an attack.

Agreed! I had to do a double take and check the URL. I thought I am reading a press release rather than actual reporting.

https://news.ycombinator.com/item?id=13155538

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#486
post #472

There's two basic kinds of distillation: 1) the massive [and dumb] method where you ask a question and use the answer as reinforcement (Black Box), and 2) more targeted distillation where you use one model to directly inform/train/guide another model (RLAIF). The latter is basically fine-tuning the model with direction from another model. Thousands of businesses do this every day to fine-tune. This is almost certainl…

Stupid question: I was under the impression that these models were trained on PB of data. Surely the amount of questions/response they can extract from querying a bigger model (Claude) is fairly modest. How is it not a drop vs the training dataset?

This might be like an observational study vs a study with a control?

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#487
post #431
post #408

Hypocrisy is a form of corruption. Anthropic's IP was created by harvesting and "distilling" other people's IP. Copyrighted materials, and the commons... which they have essentially privatized. The commercial goal is to avoid competition. One of the main worries for AI is "commoditization" which has come to mean "not a monopoly." To that end, it doesn't matter is the competitor is Chinese American or other. Their mot…

> Anthropic's IP was created by harvesting and "distilling" other people's IP. Copyrighted materials, and the commons... which they have essentially privatized. Anthropic and others argue that because LLMs don’t output full copyrighted works word for word - hence their LLMs aren’t infringing on copyright laws. I think (if this ever comes to that) Chinese lab should use same arguments against Anthropic. UPDATE: this i…

> Anthropic and others argue that because LLMs don’t output full copyrighted works word for word - hence their LLMs aren’t infringing on copyright laws.

That surely can't be what they argue, because I'm sure I can't translate a copyrighted book into a different language and say "that's fine, it's not word-for-word".

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#488

How can there be any moat for AI ever, if you can just steal a model by talking to it?

This is what I find the most fascinating about the people arguing that you can copyright-wash anything (e.g. FOSS code) by passing it through an LLM. Surely that same logic applies to the LLM itself?!

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#490
post #340
post #246

Earlier quoted context omitted.

The US labs do seem to have announced a lot of licensing deals though, and are buying things today due to the previous lawsuits. At what point will we be better to support a lab that pays (some) licenses today vs the ones that pay none? Some of the deals are in the hundreds of millions, so I suspect licensing is over a billion today? (Pure guess). That might become a big disadvantage in a price (or content) war.

At the very least the public should receive full open-weight open-source models in return for their transgressions. Failing that, may I suggest the guillotine?

In the US the courts are also pursuing labs that open their models: Meta's current court case is over the training data of the llama models they released openly.
Post reply on HN