Live data from Hacker News

Anthropic says Alibaba illicitly extracted Claude AI model capabilities

reuters.com

721–730 of 1001 posts

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#721

I'm looking forward to the trial where Anthropic will have to disclose sources of their training data, and then explain why they are entitled to charging customers for using regurgitated training data but Alibaba which trains their models on Anthropic's models are not. Should be fun. Edit: clarification

They already did and paid 1.5B https://authorsguild.org/advocacy/artificial-intelligence/wh...

That's only a fraction of the training data.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#722

Distillation is fundamentally impossible to protect against. All you can do is slow them down. Change my view. Eventually these Chinese companies will release some extension like Honey, which will sit on top real, non-Chinese clients and send everything to China anyway. It's over.

Jensen Huang likely agreed with you and tried to change Dario Amodei's view on that, but that attempt appeared to have failed.

So there's that.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#723

This is great for competition! Chinese vendors offering a cheaper solution = what economics told me the free market was all about. I also learnt that Anthropic should get better at what they do if they want to compete. If not, somebody else will win. Or does this not apply to huge US corporations any more?

China aren't offering a cheaper solution. They are subsidizing an existing one (which is already subsidized) in order to gain foothold. The difference is that in the US subsidies come from VC, while OP implies subsidies come from the AI labs that buy the training data (which may as well also be VC backed, so just one extra hop). This isn't "the market working as intended", this is an exhaustion fight to the bottom wh…

doesnt VC money subsidise stuff all the time? Isnt that how Uber and AirBNB undercut competition?

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#724

Here's what is happening: Chinese resellers are offering Claude tokens at 70-90% below official Anthropic API prices. They achieve this by reselling capacity from pooled Claude Max accounts, payments fraud, and also reselling the model output & reasoning chains to various Chinese labs. They are subsidizing model access in exchange for user logs and reasoning traces, which they then sell as training data, allowing the…

Considering how Claude and GPT were trained selling this as training data is completely justified.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#725
post #322

Earlier quoted context omitted.

https://research.nvidia.com/labs/lpr/slm-agents/ - Distillation data is a natural byproduct of using these models. There's no effective defence against it. Anthropic is degrading thinking blocks to summaries to slow it down and hide model internals, but in the end, the math says you're SOL and it works on MNC/Large Corporate scale well enough that the moment cost becomes a priority, you're left without the lock in yo…

Byproduct? It’s essentially the only part of an LLM that is useful, because it’s the WHOLE product! It’s the same reason why DRM for audio and video is a non sequitur - if you want a person to see or hear audio or video, eventually at the end of the chain, it’s going to be converted to audio for the ear and light for the eyes - that’s why you attach your tap. Without a model generating tokens, what’s the point. So if…

That's why the harness is moving server-side: because generating tokens is not the actual point of the model, not for the users. Especially with tool calling giving us agents that can act, most of the tokens generated are not, themselves, critical to the end users. Specifically, a lot of tokens goes into orchestrating actual tool calls, and then most "thinking tokens" are only relevant to users only in so far as they help users keep track of and verify what the LLM is doing. So all those tokens can be hidden or replaced by partial summaries, and all of that can happen server-side, and then there's very little to distill from.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#727

Reminds me a bit of the anecdote of Steve Jobs complaining about people ripping off the Mac GUI, in the mid to late 1980s, when he gave no public acknowledgement to the work done by Xerox on the Alto and Star operating system. "you're trying to rip off what I've already ripped off!" Crawl the whole Internet to build a gargantuan sized LLM and then complain you're being copied...

Not just the whole internet, but commit commercial copyright infringement and settle class action out of court with authors whose books you pirated.

https://www.authorsalliance.org/2025/09/07/the-anthropic-set...

"One rule for thee, a different rule for me." - Dario

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#728

Here's what is happening: Chinese resellers are offering Claude tokens at 70-90% below official Anthropic API prices. They achieve this by reselling capacity from pooled Claude Max accounts, payments fraud, and also reselling the model output & reasoning chains to various Chinese labs. They are subsidizing model access in exchange for user logs and reasoning traces, which they then sell as training data, allowing the…

Good for them!
Post reply on HN