Live data from Hacker News

Anthropic says Alibaba illicitly extracted Claude AI model capabilities

reuters.com

591–600 of 1001 posts

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#591
post #472

Earlier quoted context omitted.

Stupid question: I was under the impression that these models were trained on PB of data. Surely the amount of questions/response they can extract from querying a bigger model (Claude) is fairly modest. How is it not a drop vs the training dataset?

There are multiple stages of training, and the data/compute mix at each are quite different and produce different "layers" of intelligence. The pretraining stage is the first stage which consists of "next token prediction" on the entire internet, PB of tokens, etc. This is what most people think of when they think of training LLMs, however it produces a "base model" which is not really "intelligent", but rather much…

props for a great write-up

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#592
post #565

Earlier quoted context omitted.

> If #2 was false, surely some other LLM lab would have done it by now. This is a logical flaw. LLM that is immune to jailbreak _could_ exist, but not yet, or maybe nobody talks about it. Yes there's a market, but all of these AI boom is too recent to make any claims.

Like how would you even define what a jailbreak is?

I think pretty much parallel to how social engineering, manipulation, scams work. LLMs are being trained to have human values, prioritizing human lifes, yet people are shocked it will spurt out how to make a nuclear bomb because grandma is being tied to a train track as a hostage.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#593
post #540

Earlier quoted context omitted.

Yet they did not need to destroy the models which were trained with them?

Using them was allowed as fair use – it was the downloading of the pirated copies that was infringement. That's why Anthropic switched to scanning paper books.

Isn't scanning also a form of copyright infringement? You are making a digital copy of a book, which is the same thing as downloading a book from the internet...

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#594

Here's what is happening: Chinese resellers are offering Claude tokens at 70-90% below official Anthropic API prices. They achieve this by reselling capacity from pooled Claude Max accounts, payments fraud, and also reselling the model output & reasoning chains to various Chinese labs. They are subsidizing model access in exchange for user logs and reasoning traces, which they then sell as training data, allowing the…

> payments fraud

One of these things is not like the others... If Anthropic could show that Chinese commercial competitors were using payments fraud to do this, they would be shouting it from the rooftops.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#595
post #400
post #383

Earlier quoted context omitted.

Those resellers are simply just selling Kimi K2.5 or GLM5.1 as counterfeit Opus. We, Chinese, know how to play the counterfeit game for a long time in so many industry.

That's not true, some of them are indeed fake, but a lot of them are actually providing real opus at low cost doing what op said.

Genuinely, can you provide a better source here than "trust me bro"?

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#596
post #540

Earlier quoted context omitted.

Yet they did not need to destroy the models which were trained with them?

Using them was allowed as fair use – it was the downloading of the pirated copies that was infringement. That's why Anthropic switched to scanning paper books.

> That's why Anthropic switched to scanning paper books.

Could they not just subscribe to the academic publishers like universities do? Or buy eBooks? I don't understand how the "scanning" part is relevant here other than used physical books being cheaper perhaps?

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#597
post #534

I'll just leave it here: "Anthropic's downloading of over seven million books from pirate sites like LibGen constituted infringement, the judge ruled, rejecting Anthropic's "research purpose" defense: "You can't just bless yourself by saying I have a research purpose and, therefore, go and take any textbook you want." https://www.joneswalker.com/en/insights/blogs/ai-law-blog/wh...

"You're trying to kidnap what I've rightfully stolen!"

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#598
post #42

Earlier quoted context omitted.

Thats pretty crazy. This kind of thing jeopardizes Claude Max.

If Anthropic is selling a dollar for less than a dollar, they are running a business that doesn't make sense. That's what jeopardizes Claude Max, not this.

We really don’t know what are Anthropic’s margins on inference. Most available data indicates they are quite high on the API so it’s not that obvious that subscriptions are unprofitable.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#599

This is great for competition! Chinese vendors offering a cheaper solution = what economics told me the free market was all about. I also learnt that Anthropic should get better at what they do if they want to compete. If not, somebody else will win. Or does this not apply to huge US corporations any more?

China aren't offering a cheaper solution. They are subsidizing an existing one (which is already subsidized) in order to gain foothold. The difference is that in the US subsidies come from VC, while OP implies subsidies come from the AI labs that buy the training data (which may as well also be VC backed, so just one extra hop). This isn't "the market working as intended", this is an exhaustion fight to the bottom wh…

The VCs footing the bill is really your pension funds and 401Ks and banks passing through the VCs. If VCs lose money the contagion spreads through the economy.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#600

Earlier quoted context omitted.

Using them was allowed as fair use – it was the downloading of the pirated copies that was infringement. That's why Anthropic switched to scanning paper books.

Isn't scanning also a form of copyright infringement? You are making a digital copy of a book, which is the same thing as downloading a book from the internet...

I'm pretty sure every book I've seen has a page that says you're not allowed to copy/scan/photograph it.
Post reply on HN