Live data from Hacker News

Anthropic says Alibaba illicitly extracted Claude AI model capabilities

reuters.com

841–850 of 1001 posts

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#841
This is why I don’t understand the concerns about “our AI overlords” monopolizing all the gains from AI. It doesn’t seem like there’s much of a moat around the models themselves. So the race is mainly about compute. But compute is subject to power law effects. I remember Intel building the first Teraflop computer (ASCI red) in 1996. It was the size of a house. By 2014 you had more compute and 50% more memory in an off the shelf dual processor server system.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#842

“Distillation attack” are we joking here. If anything these models should be compelled to be public since they have been trained off public data. What an absurd overreach to call this an attack. It’s clear they are scapegoating national security and China at this point to build an anti-competitive moat. I generally really like Anthropic’s work and models but stuff like this scares me for the future. We are positionin…

> The public’s life is getting worse while these companies consolidate power using data they stole from the public

How can you “steal” public information?

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#843

Earlier quoted context omitted.

> If anything these models should be compelled to be public since they have been trained off public data I'm starting to come around to this idea TBH. For a while my position was: "these companies have invested billions into training these models, therefore they should be able to control them and profit off them" but looking deeper at where they got their training data, my view is starting to shift. IMHO I feel like…

labs invest multiple billion dollars a year each in private data, and that number is growing. internet training data is not where frontier capabilities come from, this view is outdated

When did they start doing so? We all know that they DID train on all the available public information, so at what point did they stop? Is the public information still in the training set? If so, they should STILL release ALL the data as public, as they are including training data that was acquired without permission.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#844

Earlier quoted context omitted.

> If anything these models should be compelled to be public since they have been trained off public data I'm starting to come around to this idea TBH. For a while my position was: "these companies have invested billions into training these models, therefore they should be able to control them and profit off them" but looking deeper at where they got their training data, my view is starting to shift. IMHO I feel like…

labs invest multiple billion dollars a year each in private data, and that number is growing. internet training data is not where frontier capabilities come from, this view is outdated

Does this private data come from places like Reddit, Twitter, etc., where it’s contributed by users? I think it is unethical for these companies to accept payment for user-contributed data.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#845
post #384

Earlier quoted context omitted.

Yikes, no thank you

Do you have a substantive reason why you dislike this? What is the problem if it's opt-in? Nobody is forcing you to share your usage if you don't want to. I'd prefer it if all the model builders could train on my usage rather than being limited to a single company. That'll hopefully help make all the models better in the long-term.

Very substantive - that data can be highly sensitive, and I don’t trust all model companies

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#846

Earlier quoted context omitted.

> If anything these models should be compelled to be public since they have been trained off public data I'm starting to come around to this idea TBH. For a while my position was: "these companies have invested billions into training these models, therefore they should be able to control them and profit off them" but looking deeper at where they got their training data, my view is starting to shift. IMHO I feel like…

labs invest multiple billion dollars a year each in private data, and that number is growing. internet training data is not where frontier capabilities come from, this view is outdated

No, you're talking about fine tuning and most of it is coming from your customers or someone else's. Get off ya high horse.

Copyright needs abolishing.

Companies can't be trusted with societies need for open progress.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#847
I mean I believe in protecting your company's IP, but IP and patent law is absurd these days, designed to protect investors and their fake money rather than actual inventors (who usually get no proceeds/are shafted).

They trained from the internet, so if someone trains from them it's fair game. Their clever tech should be in the mechanism with which it uses to provide an answer, not the answer itself.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#849

Earlier quoted context omitted.

> If anything these models should be compelled to be public since they have been trained off public data I'm starting to come around to this idea TBH. For a while my position was: "these companies have invested billions into training these models, therefore they should be able to control them and profit off them" but looking deeper at where they got their training data, my view is starting to shift. IMHO I feel like…

labs invest multiple billion dollars a year each in private data, and that number is growing. internet training data is not where frontier capabilities come from, this view is outdated

This is a misleading statement. The "private data" is still largely publicly produced data that has been curated through private agreements instead of scraping, such as reddit posts/comments (this is the "third-party data agreements" that companies like OpenAI mention). And yes, there is still a lot of processing done on this data, which is the norm for preparing training data.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#850

“Distillation attack” are we joking here. If anything these models should be compelled to be public since they have been trained off public data. What an absurd overreach to call this an attack. It’s clear they are scapegoating national security and China at this point to build an anti-competitive moat. I generally really like Anthropic’s work and models but stuff like this scares me for the future. We are positionin…

> If anything these models should be compelled to be public since they have been trained off public data. Isn't that a bit like saying if you read books in a public library to pick up a new skill you should work for free? > What an absurd overreach to call this an attack. Would it be an attack to take your meal by force if you used a public recipe to prepare the meal?

> Isn't that a bit like saying if you read books in a public library to pick up a new skill you should work for free?

Only if you’re trying to muddy the waters. No, obviously it’s not. One can also support licensing for driving a car on public roads but not for walking, even though both involve traveling. This is only confusing to people pretending to be confused, for effect.

> Would it be an attack to take your meal by force if you used a public recipe to prepare the meal?

“You wouldn’t download a car…” (unless it worked like copying an MP3, then, of course, you would, everyone would)

It’s as if you’re using terrible analogies and comparisons because stronger ones don’t exist. Great news for the AI-should-be-open crowd.

Post reply on HN