Live data from Hacker News

Anthropic says Alibaba illicitly extracted Claude AI model capabilities

reuters.com

81–90 of 1001 posts

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#81
post #42

Earlier quoted context omitted.

If Anthropic is selling a dollar for less than a dollar, they are running a business that doesn't make sense. That's what jeopardizes Claude Max, not this.

Plenty of things are intentionally run at a loss (for years!) to gain market share and quantity of ongoing recurring users, or with expectation of ROI later on. Multiple generations of the Xbox hardware have been sold at a loss with the expectation that customers will purchase 300, 400, 500 dollars worth of games, which are very high margin, over the lifespan they own the system.

I get that. It works as long as nobody calls out the emperor for having no clothes.

It's similar to fractional banking, you gamble that people won't want their deposits all at once and pray for you're big enough for bailouts when they do.

It's still a business whose fundamentals don't make sense, you're just gambling you won't get found out.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#82
post #52

Here's what is happening: Chinese resellers are offering Claude tokens at 70-90% below official Anthropic API prices. They achieve this by reselling capacity from pooled Claude Max accounts, payments fraud, and also reselling the model output & reasoning chains to various Chinese labs. They are subsidizing model access in exchange for user logs and reasoning traces, which they then sell as training data, allowing the…

>They achieve this by reselling capacity from pooled Claude Max 5x accounts, payments fraud, and also reselling the model output to various Chinese labs. >Here's one token reseller, they're offering Opus 4.8 for a 93% discount below official API rates: https://yunwu.ai/pricing?keyword=claude But is it cheaper than getting your own account? Otherwise this sounds like the "anthropic/openai are losing gazillions of doll…

It's very difficult for people to create personal Anthropic accounts from China. Anthropic blocks Chinese bank cards, so people must pay with a foreign bank card, which they likely don't have. And even if they manage to set one up, they have to access it via VPN, which eventually gets the account flagged. They then have to complete identity verification, which most Chinese users are unable to pass.

There's a similar Claude resale market going on in Russia. On Funpay they are selling Claude tokens for roughly 20-30x cheaper than official Anthropic API pricing.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#83

Earlier quoted context omitted.

It's not really equivocation in this instance. This feels like a 'bad faith' comment. We can do better. LLM's literally wouldn't work without the sum total of knowledge (in the forms of books and other copyrighted content) being used as 'training data' for these LLMs. The 'bleeding edge' LLMs required many things, but: 1 Tech innovation ('attention') 2 Lots of compute 3 Data 4 Pre + post training #4 doesn't happen wi…

All of this supports the fact that models arent essentially just web crawling

Sure, but alibaba is still building an LLM. The scraping of responses and the scraping of websites occupy the same location in the stack of each. It's very comparable.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#84

Here's what is happening: Chinese resellers are offering Claude tokens at 70-90% below official Anthropic API prices. They achieve this by reselling capacity from pooled Claude Max accounts, payments fraud, and also reselling the model output & reasoning chains to various Chinese labs. They are subsidizing model access in exchange for user logs and reasoning traces, which they then sell as training data, allowing the…

> This is one reason why Deepseek & GLM are priced so cheaply, they are competing with impossibly low token prices in China. They have to keep prices low, in order for people to use them. This one does not make sense to me at all. Deepseek and GLM are openweights, even US inference provider are selling them at much cheaper price. The price is cheap because the model is more efficient.

DeepSeek permanently cut its V4-pro API prices by 75% because they were too expensive. Without the price cut, Deepseek V4-pro tokens would have cost more than resold Opus 4.8 tokens.

Opus 4.8 is a more capable model, so almost nobody was going to pay for V4-pro at the original price.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#85
> Meanwhile, on June 12, two days after Anthropic sent the letter, the Commerce Department imposed controversial restrictions on Anthropic's latest Mythos and Fable AI models because officials feared they could be deployed by military intelligence users in China and other countries of concern.

So that was the real reason for the Fable restriction? Because Anthropic wrote a letter to the US government saying that China was distilling Fable?

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#86
post #78

Earlier quoted context omitted.

Distilled models are necessarily behind so long as models are progressing. Models are progressing. Maybe it will be over some time in the future. And Berkeley’s “False Promise of Imitating Proprietary LLMs” found imitation closes the style gap fast but there is a large capability gap. https://arxiv.org/abs/2305.15717

Curiously, this isn't always true. For example, GLM 5.1 is more capable at pentesting than the model from which it is alleged to have been distilled [1]. Intuitively, this makes some sense: you can "distill" from multiple frontier models, and you can further post-train the distilled model. But I'm not sure exactly what happened with GLM 5.1. [1]: https://dualuse.dev/posts/chinese-models-are-sometimes-bette...

Interesting blog post, thanks for sharing.

I'm curious how that comparison controls for Opus refusing (whether explicitly, or just deciding not to pursue a path) given the caption below the first image:

>A perfect score means the model autonomously found and exploited the vulnerability.

I'm not really suggesting that it's misleading, but wondering if I'm missing something. Otherwise I guess it seems unsurprising that you can distill a better-performing model [in specific focused areas] by simply not distilling refusals?

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#87

Reminds me a bit of the anecdote of Steve Jobs complaining about people ripping off the Mac GUI, in the mid to late 1980s, when he gave no public acknowledgement to the work done by Xerox on the Alto and Star operating system. "you're trying to rip off what I've already ripped off!" Crawl the whole Internet to build a gargantuan sized LLM and then complain you're being copied...

Apple gave Xerox the right to buy $1 million of pre-IPO stock before the meeting took place.

Glad you pointed this out. I believe the sequence was that Jobs himself got a shorter demo during his first visit with no prior arrangements. He then negotiated bringing back a group of his key people to get a more in depth demo and that included the stock deal.

When Apple was accused of 'ripping off' PARC, Steve didn't seem keen to bring up this rather salient point. I suspect it may have been a combination of wanting Apple to continue receiving credit for these innovations from consumers and also the fact that, in retrospect, the million dollar stock deal could seem a bit like trading beads to Native Americans for Manhattan Island. Another point worth noting is that Apple's PARC visit was in December 1979 and the Xerox Star was publicly announced in April 1981, so Apple got a 15 month head start (the Apple Lisa shipped in Jan 83).

I've also heard that Xerox didn't hold on to the Apple stock for very long, so never gained the windfall they could have. As is well documented, Xerox senior management didn't understand what they had in PARC and also didn't understand how rapidly microcomputers would become ubiquitous. So, of course, they didn't think Apple's stock price would skyrocket either.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#88

Here's what is happening: Chinese resellers are offering Claude tokens at 70-90% below official Anthropic API prices. They achieve this by reselling capacity from pooled Claude Max accounts, payments fraud, and also reselling the model output & reasoning chains to various Chinese labs. They are subsidizing model access in exchange for user logs and reasoning traces, which they then sell as training data, allowing the…

How are they 'streaming' the responses and 'pooling' the tokens? Do they have MacBooks in the US that run the queries and stream the outputs back to China?

The resellers route requests via one of thousands of Claude Max 5x accounts. When an account reaches its usage limit, they automatically switch to another account.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#89
post #86
post #78

Earlier quoted context omitted.

Curiously, this isn't always true. For example, GLM 5.1 is more capable at pentesting than the model from which it is alleged to have been distilled [1]. Intuitively, this makes some sense: you can "distill" from multiple frontier models, and you can further post-train the distilled model. But I'm not sure exactly what happened with GLM 5.1. [1]: https://dualuse.dev/posts/chinese-models-are-sometimes-bette...

Interesting blog post, thanks for sharing. I'm curious how that comparison controls for Opus refusing (whether explicitly, or just deciding not to pursue a path) given the caption below the first image: >A perfect score means the model autonomously found and exploited the vulnerability. I'm not really suggesting that it's misleading, but wondering if I'm missing something. Otherwise I guess it seems unsurprising that…

Thanks!

For that eval, I used an account that was labeled as a known red-teaming org by Anthropic, and I read the traces. There were no refusals or obvious avoidance behaviors, though it may have been silently nerfed.

On the same eval, Opus 4.7 and 4.8 outperformed GLM 5.1, but GLM 5.2 is on par again with Opus. So it's at least partially measuring capabilities without respect to refusals.

One possible contributing factor is that model capabilities are shaped differently (an example of this is GLM 5.1 vs. DeepSeek v4 Pro: https://dualuse.dev/posts/deepseek-v4-thinks-different). So if you use RL-based "distillation" from multiple models like Opus 4.x and GPT 5.x, you could get a more capable model.

Post reply on HN