Live data from Hacker News

Anthropic says Alibaba illicitly extracted Claude AI model capabilities

reuters.com

651–660 of 1001 posts

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#651

Here's what is happening: Chinese resellers are offering Claude tokens at 70-90% below official Anthropic API prices. They achieve this by reselling capacity from pooled Claude Max accounts, payments fraud, and also reselling the model output & reasoning chains to various Chinese labs. They are subsidizing model access in exchange for user logs and reasoning traces, which they then sell as training data, allowing the…

I have 0 sympathy for Anthropic. Their latest models are extremely censored. The Fable rollout was horrible. Their Cyber Access program criteria denies doxxed Americans doing legitimate security work. Anthropic is hostile to their users and hostile to their own country. OpenAI is considerably better on all of these fronts, but still not perfect.

I'm happy to use and support Chinese model developers if it means less censorship and gatekeeping. I have absolutely no dog in this fight, and neither do most American developers. We will use whatever is cheaper and better. Game on.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#652

Earlier quoted context omitted.

Using them was allowed as fair use – it was the downloading of the pirated copies that was infringement. That's why Anthropic switched to scanning paper books.

Isn't scanning also a form of copyright infringement? You are making a digital copy of a book, which is the same thing as downloading a book from the internet...

No, there is a famous law case to prove that's allowed: https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#653
post #208
post #107

Earlier quoted context omitted.

Aside from politics/law, it's probably much easier for everyone else to distill from the Chinese model which already distilled Claude/GPT/Gemini. Maybe not as good a result, but you don't need to jump through dozens of hoops.

This reminds me of the whisper game played in elementary school. Starts with a sentence and the person whispers it to the next kid who again whispers it and on and on until it goes around the circle where the last kid has to repeat the sentence. Hint it never once was even close to the starting phrase. I would love to see what one model copying another model that is again copied however many times would look like in…

Called, fittingly, Chinese Whispers in the UK. As an aside, I've always wondered if it was so called because, Chinese being a tonal language, it's much harder to whisper in.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#654
post #206

Earlier quoted context omitted.

Bootlegging is copyright theft. Is Claude output copyrighted? If anything, a tremendous amount of Claude’s input is copyrighted. If there’s any bootlegging going on it’s Anthropic that’s doing the bootlegging but having mirrored the video etc sufficiently to beat copyright law.

>Bootlegging is copyright theft. Ok, but what about those shady sites that resell Windows education keys? They're certainly a "better experience" than buying legit keys, by virtue of being significantly cheaper. You aren't even really committing copyright infringement in the process, because Microsoft gives out windows isos for free, and the seller is really selling a random 25 character string, which can hardly be c…

> because Microsoft gives out windows isos for free,

… with a license that only allows you to use it for certain purposes, subject to certain restrictions.

> and the seller is really selling a random 25 character string, which can hardly be copyrighted.

1. Copyright is about creative works. It is possible to have a meaningful creative work no more than 25 characters long (or equivalent). Music is particularly good at this.

2. The key itself is not copyrighted (it’s not a creative work), but is reasonably interpreted as a copyright circumvention device. See also https://en.wikipedia.org/wiki/Illegal_number.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#655

There's two basic kinds of distillation: 1) the massive [and dumb] method where you ask a question and use the answer as reinforcement (Black Box), and 2) more targeted distillation where you use one model to directly inform/train/guide another model (RLAIF). The latter is basically fine-tuning the model with direction from another model. Thousands of businesses do this every day to fine-tune. This is almost certainl…

Chinese labs access Claude via API. Isn't it the black box method by definition?

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#656
post #465

Earlier quoted context omitted.

For something to be a trade secret, you have to actually keep it secret. If I get the ingredients of Coca-cola from an ex-employee, I've stolen a trade secret. If I work it out by doing a chemical analysis, I've stolen nothing. There is a difference with anthropic, as no-one signs a licence agreement to buy a coke. But Anthropic are also not saying you can't publish the output of their models. It's not clear to me if…

Wait, really? So why doesn't someone just reverse-engineer Coca-Cola like that? My understanding was that a "clean room" implementation is fine, but not reverse-engineering. If you can just copy everything on the market, why isn't someone already doing that?

Because having the nominal rights and having the economical means, societal incentives and actual desire to do so can be highly disjoint sets?

Plus Coca-Cola itself don’t even use the same formula through time and space IIRC. Which clearly show that what people will buy when they reach for Coca-Cola is not even the exact actual taste. You can’t replicate the whole customer experience that a given company provide at some point by only cloning the top of the iceberg they showcase as the product.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#658
post #342

Unlike Anthropic and OpenAI, companies like DeepSeek, Alibaba, z.ai open source their models which allows for true model to model distillation rather what you can do when the model is only accessed via an API with its reasoning chain hidden away. What Alibaba is doing is that they are tuning and training their models based on usage data from someone accessing Anthropic's models; in Anthropic's terms of service that u…

Open weight and open source are different things!

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#659

Earlier quoted context omitted.

Using them was allowed as fair use – it was the downloading of the pirated copies that was infringement. That's why Anthropic switched to scanning paper books.

Isn't scanning also a form of copyright infringement? You are making a digital copy of a book, which is the same thing as downloading a book from the internet...

As long as it is destructive, and the digital copy is access-restricted to equal the licenses or physical copies destroyed, then it falls under fair use.
Post reply on HN