Live data from Hacker News

Anthropic says Alibaba illicitly extracted Claude AI model capabilities

reuters.com

851–860 of 1001 posts

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#851

“Distillation attack” are we joking here. If anything these models should be compelled to be public since they have been trained off public data. What an absurd overreach to call this an attack. It’s clear they are scapegoating national security and China at this point to build an anti-competitive moat. I generally really like Anthropic’s work and models but stuff like this scares me for the future. We are positionin…

> The public’s life is getting worse while these companies consolidate power using data they stole from the public How can you “steal” public information?

really? You know this just like everyone else: Just because the information is available publicly, does not mean that you can do whatever you want with the information. Copyright exists for a reason, and if the copyright lobby is going to continue to push for the poor poor media companies to keep their copyrights, then we should do the same towards the AI companies. So yes, they Stole the information from everyone else, and they keep doing so, as you can see their scanners still hitting every website on the web to get an updated dataset. It does not matter what they do AFTER they steal all the information, as they already stole it.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#853

Earlier quoted context omitted.

> If anything these models should be compelled to be public since they have been trained off public data I'm starting to come around to this idea TBH. For a while my position was: "these companies have invested billions into training these models, therefore they should be able to control them and profit off them" but looking deeper at where they got their training data, my view is starting to shift. IMHO I feel like…

labs invest multiple billion dollars a year each in private data, and that number is growing. internet training data is not where frontier capabilities come from, this view is outdated

Great way to launder illegally obtained data too.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#855

“Distillation attack” are we joking here. If anything these models should be compelled to be public since they have been trained off public data. What an absurd overreach to call this an attack. It’s clear they are scapegoating national security and China at this point to build an anti-competitive moat. I generally really like Anthropic’s work and models but stuff like this scares me for the future. We are positionin…

Should Google search index be forced to be public too?

Honestly, yes it should in some form. If their index contains the actual data from the sites, and they are making that information public in one way or another, then it should be available as a downloadable dataset.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#856

“Distillation attack” are we joking here. If anything these models should be compelled to be public since they have been trained off public data. What an absurd overreach to call this an attack. It’s clear they are scapegoating national security and China at this point to build an anti-competitive moat. I generally really like Anthropic’s work and models but stuff like this scares me for the future. We are positionin…

> If anything these models should be compelled to be public since they have been trained off public data I'm starting to come around to this idea TBH. For a while my position was: "these companies have invested billions into training these models, therefore they should be able to control them and profit off them" but looking deeper at where they got their training data, my view is starting to shift. IMHO I feel like…

I'm not taking sides here but this situation is not so black and white and it has always been the darker side of capitalism.

The concept of Intellectual property exists not because it's fair but because it creates incentive to make said "intellectual property" exist. If intellectual property can be instantly copied by a competitor... why would I spend a dime to even create such a thing? I want to profit off of what I make because I'm a capitalist and money is what drives me (as a capitalist).

Anthropic models wouldn't exist if they couldn't keep a unholy grip on it. Same with openAI. Same with many life saving drugs.

Of course everyone here is talking about the obvious stuff like how it's morally wrong to with-hold life saving drugs or to have AI literally take over the world and be under the control of one company and all of this is true. But it is also true that greed is the engine that drives our economy and if you want our economy to produce "intellectual property" you must allow people to "capitalize" on that greed.

There are two controversial issues here. What is moral/fair? And what is realistically practical in optimizing the economy if said economy is based on money.

The distillation in my mind is a win for practicality because Competition also drives our economic engine. First you don't want a monopoly, but you also don't want these models to be so damn open that there's zero incentive to make them.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#857

Earlier quoted context omitted.

> If anything these models should be compelled to be public since they have been trained off public data. Isn't that a bit like saying if you read books in a public library to pick up a new skill you should work for free? > What an absurd overreach to call this an attack. Would it be an attack to take your meal by force if you used a public recipe to prepare the meal?

> Isn't that a bit like saying if you read books in a public library to pick up a new skill you should work for free? Only if you’re trying to muddy the waters. No, obviously it’s not. One can also support licensing for driving a car on public roads but not for walking, even though both involve traveling. This is only confusing to people pretending to be confused, for effect. > Would it be an attack to take your meal…

I think the analogies are appropriate. Anthropic took public data and added value on top of it. It is that added value that Alibaba is targeting. If it was the underlying data, that's freely available.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#858
post #274

Earlier quoted context omitted.

Externally subsidized predatory pricing is the opposite of a free market.

So all those companies selling at a loss to gain market share aren't part of the free market? Like openai, anthropic, and SpaceX?

Externally subsidized predatory pricing is the opposite of a free market — precisely because it sells things at below market rates.

Free markets are where players compete on quality, efficiency, and supply. Prices are a result of cost and supply and provide real information on these factors. Competition for customers selects the most effective and efficient producer.

Sustained efforts of selling at a loss to gain market share is the exact opposite. The entire purpose is to corrupt the free market by sending false price signals which SUPPRESS free market competition and push market share to whoever can burn the most capital (whilst providing an actual service/product), not whoever is most efficient or highest quality or lowest actual price provider.

Uber and AirBnB are better examples of your "selling at a loss to gain market share", where they burned capital to undercut prices for close to a decade on falsely low pricing to destroy incumbents.

Spending on R&D while developing expensive technology is different and arguably very much a part of a free market, and is not what I was talking about.

Spending capital to steal your competitors' technology, and then spending more of it to make it available at below-market rates, is absolutely not a free-market activity.

Just because it is not stopped by someone enforcing a free market, does not make it a free market.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#859

There's two basic kinds of distillation: 1) the massive [and dumb] method where you ask a question and use the answer as reinforcement (Black Box), and 2) more targeted distillation where you use one model to directly inform/train/guide another model (RLAIF). The latter is basically fine-tuning the model with direction from another model. Thousands of businesses do this every day to fine-tune. This is almost certainl…

Chinese companies are engaging in anti-competitive practices, as usual. They are rogue actors on the economic scene. If it were feasible, they'd be widely banned, and for good reason.

Bringing more competition is "anti-competitive" now.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#860

Earlier quoted context omitted.

> If anything these models should be compelled to be public since they have been trained off public data I'm starting to come around to this idea TBH. For a while my position was: "these companies have invested billions into training these models, therefore they should be able to control them and profit off them" but looking deeper at where they got their training data, my view is starting to shift. IMHO I feel like…

labs invest multiple billion dollars a year each in private data, and that number is growing. internet training data is not where frontier capabilities come from, this view is outdated

> internet training data is not where frontier capabilities come from

In that case, it should be no problem for the labs to train their new models without using public data, right?

Post reply on HN