Live data from Hacker News

Anthropic says Alibaba illicitly extracted Claude AI model capabilities

reuters.com

631–640 of 1001 posts

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#631

I'm looking forward to the trial where Anthropic will have to disclose sources of their training data, and then explain why they are entitled to charging customers for using regurgitated training data but Alibaba which trains their models on Anthropic's models are not. Should be fun. Edit: clarification

And if it includes at least one GPL source, they should release the weights on GPL license.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#632

Earlier quoted context omitted.

Using them was allowed as fair use – it was the downloading of the pirated copies that was infringement. That's why Anthropic switched to scanning paper books.

Isn't scanning also a form of copyright infringement? You are making a digital copy of a book, which is the same thing as downloading a book from the internet...

Here we have a 15% limit on scanning for fair use

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#634

Earlier quoted context omitted.

It's not about how big your dataset is - it's about how you use it. I jest, but I'm also completely serious. 1T tokens from Claude can teach a model something 1T tokens scraped from the open web can't. Things like "how an LLM can problem solve effectively", or "how an LLM should use tools", or "how to construct reasoning chains", or "when to double check", or "what innate capabilities an LLM can or can't rely on". Th…

Unremarkable base model will remain an unremarkable fine-tuned model that memorised a couple thousand of input-output pairings.

Yes, neural networks are famously poor at generalising.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#635
post #538

Earlier quoted context omitted.

It seems more like the Chinese companies ar playing the dirty game, distilling through bot accounts, not letting real competition across their firewall.

So you are believing Anthropic's claim here, and it's not as if Anthropic didn't steal the data to train the model in the first place. I think the original sin doesn't give them any ability to complain. - https://www.theguardian.com/technology/2025/sep/05/anthropic...

As far as I am concerned, this is a national security matter.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#636
post #512

Earlier quoted context omitted.

> But if you show them a jailbreak of their model that bypasses their safety, they'll tell you that any model can eventually be jailbroken so don't worry about safety. They claim two things: 1) The specific, available jailbreak for Fable 5 is not dangerous - this has been confirmed by multiple experts, and there is no credible evidence against this claim (in other words, Anthropic is probably correct) 2) It is imposs…

I'm pretty sure that Gödel incompleteness theorem and its consequences pretty much guarantee #2

No actually I don't think it does and I don't think they're related.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#637

There's two basic kinds of distillation: 1) the massive [and dumb] method where you ask a question and use the answer as reinforcement (Black Box), and 2) more targeted distillation where you use one model to directly inform/train/guide another model (RLAIF). The latter is basically fine-tuning the model with direction from another model. Thousands of businesses do this every day to fine-tune. This is almost certainl…

The compute deficit of Chinese Ai companies is real, and it IS THE ONLY competitive advantage that Western companies have. The only way the U.S. keeps that edge is to prevent distillation. The only way Chinese companies can make up for the deficit in compute is to distill. There innovation in great supply on every side of the Ocean. Its about the chips. And in terms of national security, for the U.S., and for China,…

Define compute deficit?

They've been bringing out open weight models competitive with frontier models. How could they do that if they had a compute deficit?

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#638

Earlier quoted context omitted.

That sounds like it would actually be fraud.

Not if you simply say in the terms of service that it's allowed. Then suddenly it's normal (every company does this). Similarly to how the terms of service can simply say you're not allowed to sue.

It's maybe how it works in the US, but that sure as hell not how it works in the EU (my snark wants to continue as "or any sane legislature")

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#639

There's two basic kinds of distillation: 1) the massive [and dumb] method where you ask a question and use the answer as reinforcement (Black Box), and 2) more targeted distillation where you use one model to directly inform/train/guide another model (RLAIF). The latter is basically fine-tuning the model with direction from another model. Thousands of businesses do this every day to fine-tune. This is almost certainl…

The compute deficit of Chinese Ai companies is real, and it IS THE ONLY competitive advantage that Western companies have. The only way the U.S. keeps that edge is to prevent distillation. The only way Chinese companies can make up for the deficit in compute is to distill. There innovation in great supply on every side of the Ocean. Its about the chips. And in terms of national security, for the U.S., and for China,…

If saying “plz don’t distill me” is your moat, you don’t have a moat.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#640
post #570

Everyone here praising these Chinese companies for their smarts (sure they are smart) has been ignoring this very big fact, they're improvements have mostly been by being parasitic on the leading edge SOTA models, not from some inherent innovation advantage. They are as innovative as their western counterparts, but they lack the compute, so their keeping up within months of those SOTA models depends on other means, l…

>like distillation attacks. I don't blame them; its the obvious only strategy when you cant compete in compute >distillation attacks are the only vector to keep up It's demonstrably wrong, they invest in architectural improvements as well, for example, DeepSeek's compressed attention. When you lack compute, you need fast training/fast inference, and distillation alone doesn't solve it. From what I understand, that ki…

I explicitly called out the fact that there is plenty of innovation, but that we see t Lots of innovation in both Chinese and U.S. labs, and I don't think that there is a co.parative difference there.
Post reply on HN