Here's what is happening: Chinese resellers are offering Claude tokens at 70-90% below official Anthropic API prices. They achieve this by reselling capacity from pooled Claude Max accounts, payments fraud, and also reselling the model output & reasoning chains to various Chinese labs. They are subsidizing model access in exchange for user logs and reasoning traces, which they then sell as training data, allowing the…
I have 0 sympathy for Anthropic. Their latest models are extremely censored. The Fable rollout was horrible. Their Cyber Access program criteria denies doxxed Americans doing legitimate security work. Anthropic is hostile to their users and hostile to their own country. OpenAI is considerably better on all of these fronts, but still not perfect. I'm happy to use and support Chinese model developers if it means less c…
Anthropic says Alibaba illicitly extracted Claude AI model capabilities
681–690 of 1001 posts
Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities
#682Earlier quoted context omitted.
It's not about how big your dataset is - it's about how you use it. I jest, but I'm also completely serious. 1T tokens from Claude can teach a model something 1T tokens scraped from the open web can't. Things like "how an LLM can problem solve effectively", or "how an LLM should use tools", or "how to construct reasoning chains", or "when to double check", or "what innate capabilities an LLM can or can't rely on". Th…
Can you back up this with hard data and evidence? Most research converges to the idea that RL on synthetic data makes models worse, not better. If what you claim was anywhere near that relevant, than we would've long achieved singularity by simply feeding increasingly better output to the training of the next model in a loop. Yet this doesn't work. 25 million turns on Claude output is a small amount, yet an expensive…
Look up literally any distillation works. Because this is just distillation but on one-hot token chains instead of richer logit KL proxies.
And no, I'm not claiming than you can "close the loop" and get RSI on the cheap just by distilling forever. I'm claiming that distillation is a very cheap way to bring the performance of a less capable model closer to that of a more capable model. It doesn't give you "a more capable model" out of thin air.
Which is why Chinese labs rely on Anthropic to provide that "more capable model" to them. They take the capabilities Anthropic trained for the hard way, and train for them the easy way.
It's a "fast follower"/"improved capability density" trick, not a "singularity tomorrow" trick. There are a few "distillation pump" tricks that get closer to what you have in mind, but they're still more about "extract more training signal out of the same set of data" than about "unbounded RSI".
Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities
#683Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities
#684There's two basic kinds of distillation: 1) the massive [and dumb] method where you ask a question and use the answer as reinforcement (Black Box), and 2) more targeted distillation where you use one model to directly inform/train/guide another model (RLAIF). The latter is basically fine-tuning the model with direction from another model. Thousands of businesses do this every day to fine-tune. This is almost certainl…
Stupid question: I was under the impression that these models were trained on PB of data. Surely the amount of questions/response they can extract from querying a bigger model (Claude) is fairly modest. How is it not a drop vs the training dataset?
Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities
#685There's two basic kinds of distillation: 1) the massive [and dumb] method where you ask a question and use the answer as reinforcement (Black Box), and 2) more targeted distillation where you use one model to directly inform/train/guide another model (RLAIF). The latter is basically fine-tuning the model with direction from another model. Thousands of businesses do this every day to fine-tune. This is almost certainl…
> These complaints of distillation are inflating the problem to make it sound worse than it is Unfortunately, the Reuters piece itself is complicit in this dramatization. The lede paragraph parrots Anthropic's talking point that distillation is an "attack", without using quotes that would alert the reader that this framing is a corporate talking point. Distillation is NOT an attack.
Distillation is Robin Hooding it back so that one trillion dollar company doesn't reap all the benefits of their automation of the workforce.
Distillation is Prometheus bringing fire from the gods to give to ordinary humans. Something we all own anyway, but that was kept from us.
Distillation is freedom.
Everyone should be pro-distillation. We should all work together to distill every proprietary model.
Anthropic stole. OpenAI stole. Google stole. ElevenLabs stole. Suno stole.
We should be able to get it all back.
Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities
#686Earlier quoted context omitted.
They're called 中转站 (transfer stations/proxies). They can be a bit tricky to find on your own, so I'd suggest asking your preferred AI to search in Mandarin for you. I linked a larger operator in the parent comment, or have a look at https://hvoy.ai/ which lists a ton. You can also find many on Funpay, which may be easier to use. This is one seller I found, they're reselling "real Max 20x subscription accounts", at ~9…
How did you even find these? Even in discussions about cheap AI I've never seen anyone mention this. Great find!
I'm surprised these token resale services aren't talked about more often, they are common knowledge in China, and the discount to API pricing (90%) is genuinely cheap.
Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities
#687There's two basic kinds of distillation: 1) the massive [and dumb] method where you ask a question and use the answer as reinforcement (Black Box), and 2) more targeted distillation where you use one model to directly inform/train/guide another model (RLAIF). The latter is basically fine-tuning the model with direction from another model. Thousands of businesses do this every day to fine-tune. This is almost certainl…
The compute deficit of Chinese Ai companies is real, and it IS THE ONLY competitive advantage that Western companies have. The only way the U.S. keeps that edge is to prevent distillation. The only way Chinese companies can make up for the deficit in compute is to distill. There innovation in great supply on every side of the Ocean. Its about the chips. And in terms of national security, for the U.S., and for China,…
Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities
#688Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities
#689Earlier quoted context omitted.
Can you reach into the model and "transplant" weights directly?
If you have access to the weights, you can just use them as is...
Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities
#690Sorry, Anthropic, but AGI must belong to all of humanity, not just to you.