Live data from Hacker News

Anthropic says Alibaba illicitly extracted Claude AI model capabilities

reuters.com

411–420 of 1001 posts

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#411

Here's what is happening: Chinese resellers are offering Claude tokens at 70-90% below official Anthropic API prices. They achieve this by reselling capacity from pooled Claude Max accounts, payments fraud, and also reselling the model output & reasoning chains to various Chinese labs. They are subsidizing model access in exchange for user logs and reasoning traces, which they then sell as training data, allowing the…

One would think Anthropic could point Mythos at this to solve the reseller problem outright: - Purchase multiple accounts via resellers - Send messages that contain a UID - Capture these in Anthropic's logs - Shut down account. Use any metadata to identify related accounts /loop

This only shuts down the account you have bought in the first, plus a few others if it is shared.

> Use any metadata to identify related accounts

How does that work? I think this is the most important part to have an impact on the „thousand“ bot accounts.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#412
post #140

Earlier quoted context omitted.

>This is great for competition! Chinese vendors offering a cheaper solution = what economics told me the free market was all about. Yeah, like all those Chinese bootleggers selling DVDs for a few dollars rather than $20. Free market! https://news.ycombinator.com/item?id=48664814

It's quite curious how multi billion dollar enterprises can't compete with a Chinese bootlegger with a big jacket, tbh. Imagine having such a warchest and being so bad at business, lol.

Bad at business? One of them has to make the thing.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#413

Earlier quoted context omitted.

I'm sure all the artists and creators they stole from had stipulations too.

Anthropic paid one billion in a copyright settlement. That's a lot of money considering they never distributed the pirated books they trained on. Nowadays they buy copies of books, train on them, and then destroy them.

>> I'm sure all the artists and creators they stole from had stipulations too.

> Anthropic paid one billion in a copyright settlement.

Because a judge determined Anthropic was engaged in piracy.

> That's a lot of money considering they never distributed the pirated books they trained on.

This is "fruit of the poisonous tree" as it were. Distributing content derived from pirated content ("pirated books they trained on") is why Anthropic had to pay what they paid.

> Nowadays they buy copies of books, train on them, and then destroy them.

There is a case one could make that this practice could be seen as unauthorized redistribution of a derivative work intended to deprive copyright holders of legitimate revenue.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#414
Relevant article - https://www.anthropic.com/news/detecting-and-preventing-dist... (3 labs generated over 16 million exchanges with Claude through approximately 24,000 fraudulent accounts). So extraction in this context is distillation.

While it is obvious to many, a modern LLM is built in roughly three stages: the foundation (pretraining) model, then SFT/supervised fine-tuning (distillation makes it easy), then the RL/RLHF stage on top (most effort-intensive). For today's reasoning models, RL/RLHF is becoming the most compute-intensive part.

Companies like Anthropic spent millions building those fine-tuning examples. A follower can shortcut that on both cost and time by distilling, and it will keep happening: every time the frontier lab climbs higher, others will find a way to shortcut the new gap. There's very little Anthropic can do beyond fraud prevention and blocking accounts that violate their terms of service.

On the policy question, I'm completely against banning Chinese models. I'm a heavy Claude Code user and I'll keep being one. But there should absolutely be price competition. China is eating the rest of the world for breakfast, lunch and dinner on manufacturing, and it did not help to ban them. Frontier pricing can't sit at 10x a capable competitor. It doesn't need to be at par either — demand is higher, and quality, trust, and fewer tokens to finish a task are worth a premium — but 4–5x is defensible.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#415

Earlier quoted context omitted.

One would think Anthropic could point Mythos at this to solve the reseller problem outright: - Purchase multiple accounts via resellers - Send messages that contain a UID - Capture these in Anthropic's logs - Shut down account. Use any metadata to identify related accounts /loop

Maybe Fable is not as capable as thought? On the one hand they talk it up as world ending and on the other hand they can't manage bot accounts on their own service. I want to hear how this can be rationalised. From the article "every layer of control frontier US AI companies have added (geoblocking, phone verification, credit card requirements, and now live biometric KYC checks) has produced a corresponding layer of…

No system is foolproof. They'd have to be willing to throw out some % of good customers along with the bots. Amazon can do that because they have a monopoly already. Anthropic can't risk it when they're trying to grab market share.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#416

There's two basic kinds of distillation: 1) the massive [and dumb] method where you ask a question and use the answer as reinforcement (Black Box), and 2) more targeted distillation where you use one model to directly inform/train/guide another model (RLAIF). The latter is basically fine-tuning the model with direction from another model. Thousands of businesses do this every day to fine-tune. This is almost certainl…

> These complaints of distillation are inflating the problem to make it sound worse than it is Unfortunately, the Reuters piece itself is complicit in this dramatization. The lede paragraph parrots Anthropic's talking point that distillation is an "attack", without using quotes that would alert the reader that this framing is a corporate talking point. Distillation is NOT an attack.

Agreed! I had to do a double take and check the URL. I thought I am reading a press release rather than actual reporting.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#417
post #408

Hypocrisy is a form of corruption. Anthropic's IP was created by harvesting and "distilling" other people's IP. Copyrighted materials, and the commons... which they have essentially privatized. The commercial goal is to avoid competition. One of the main worries for AI is "commoditization" which has come to mean "not a monopoly." To that end, it doesn't matter is the competitor is Chinese American or other. Their mot…

Bad China is stealing our stolen IP!

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#418

Earlier quoted context omitted.

"Your extremely efficient multi petabyte internet content suction machine is ripping off my extremely efficient multi petabyte internet content suction machine" Sucking down petabytes of peoples' copyrighted content that they never granted a specific license to you to use seems to be an unavoidable and default part of the process of building any huge LLM.

So why was there crawling in 1998 but no LLMs?

I am unable to comprehend the state of mind that would lead one to ask this question.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#419

There's two basic kinds of distillation: 1) the massive [and dumb] method where you ask a question and use the answer as reinforcement (Black Box), and 2) more targeted distillation where you use one model to directly inform/train/guide another model (RLAIF). The latter is basically fine-tuning the model with direction from another model. Thousands of businesses do this every day to fine-tune. This is almost certainl…

> But if you show them a jailbreak of their model that bypasses their safety, they'll tell you that any model can eventually be jailbroken so don't worry about safety.

They claim two things:

1) The specific, available jailbreak for Fable 5 is not dangerous - this has been confirmed by multiple experts, and there is no credible evidence against this claim (in other words, Anthropic is probably correct)

2) It is impossible to build an LLM that is immune to all jailbreaks. Again, there is no credible evidence against this claim, i.e. Anthropic is again entirely correct.

If #1 was false, they could just publish the details of the jailbreak - it supposedly only works on Fable 5, so there's no possible danger.

If #2 was false, surely some other LLM lab would have done it by now. Especially since a number of governments have made it clear there is a market for such a project.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#420

Someone should setup a plugin or something for Claude Code that makes it easy to log all inputs and outputs for people who are willing and interested in sharing their usage. I don't want Anthropic to be the only company that can train on my usage, I want to share my usage so it can be used for training all new models. Once you have a system for collecting all logs, you just need a place where they can be submitted. I…

Discussed building it with my friends, obviously you might share secrets and other real reasons, but if gangs of corporations are already doing it, I don't see why we shouldn't just share it amongst the crowd too.
Post reply on HN