Live data from Hacker News

Anthropic says Alibaba illicitly extracted Claude AI model capabilities

reuters.com

431–440 of 1001 posts

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#431
post #408

Hypocrisy is a form of corruption. Anthropic's IP was created by harvesting and "distilling" other people's IP. Copyrighted materials, and the commons... which they have essentially privatized. The commercial goal is to avoid competition. One of the main worries for AI is "commoditization" which has come to mean "not a monopoly." To that end, it doesn't matter is the competitor is Chinese American or other. Their mot…

> Anthropic's IP was created by harvesting and "distilling" other people's IP. Copyrighted materials, and the commons... which they have essentially privatized.

Anthropic and others argue that because LLMs don’t output full copyrighted works word for word - hence their LLMs aren’t infringing on copyright laws.

I think (if this ever comes to that) Chinese lab should use same arguments against Anthropic.

UPDATE: this is slight hyperbole of course, not worth arguing what they actually said. The point is intent and the facts - "The Big LLMs" "distilled" collective knowledge including copyrighted works at unimaginable scale, but it's all kosher and totally not piracy/copyright infringement. Though if you're teenager torrenting an mp3 - you'll get screwed.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#432

Someone should setup a plugin or something for Claude Code that makes it easy to log all inputs and outputs for people who are willing and interested in sharing their usage. I don't want Anthropic to be the only company that can train on my usage, I want to share my usage so it can be used for training all new models. Once you have a system for collecting all logs, you just need a place where they can be submitted. I…

Discussed building it with my friends, obviously you might share secrets and other real reasons, but if gangs of corporations are already doing it, I don't see why we shouldn't just share it amongst the crowd too.

Yeah I could see it being a problem if you're doing work on closed source or repos with sensitive credentials. Since my usage has all been on open source projects I'd be happy to share everything I'm doing if it can help train better models.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#433
post #246

Earlier quoted context omitted.

The US labs do seem to have announced a lot of licensing deals though, and are buying things today due to the previous lawsuits. At what point will we be better to support a lab that pays (some) licenses today vs the ones that pay none? Some of the deals are in the hundreds of millions, so I suspect licensing is over a billion today? (Pure guess). That might become a big disadvantage in a price (or content) war.

> At what point will we be better to support a lab that pays (some) licenses today vs the ones that pay none? Why is a lab that pays all licenses today not on your list? Is ethics and morality that low on your radar?

I agree that that's a more consistent position for the people criticising the data slurping. But I don't see people advocating those open-data models in these threads? It's usually about defending the zero-licensing competitors.

My (limited, outsider) understanding is that due to the court cases US labs are pressured to be legal now (for instance, bulk scanning purchased books instead of Books3, and the licensing deals with media companies). But international labs are not. The "not licensing everything" statement is more about current copyright law not requiring licensing of everything. But that question is still up in the air as cases are ongoing.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#434
post #431
post #408

Hypocrisy is a form of corruption. Anthropic's IP was created by harvesting and "distilling" other people's IP. Copyrighted materials, and the commons... which they have essentially privatized. The commercial goal is to avoid competition. One of the main worries for AI is "commoditization" which has come to mean "not a monopoly." To that end, it doesn't matter is the competitor is Chinese American or other. Their mot…

> Anthropic's IP was created by harvesting and "distilling" other people's IP. Copyrighted materials, and the commons... which they have essentially privatized. Anthropic and others argue that because LLMs don’t output full copyrighted works word for word - hence their LLMs aren’t infringing on copyright laws. I think (if this ever comes to that) Chinese lab should use same arguments against Anthropic. UPDATE: this i…

Isn’t the output of LLMs completely copyright-free in the US?

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#435

Earlier quoted context omitted.

Can you reach into the model and "transplant" weights directly?

You can do things like that - one example is averaging weights between related models - but not with Anthropic's models, because outsiders don't have access to the weights.

Weights are just data a server, so we don't know outsiders have access (either via breakin or arrangement).

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#436
post #375

Earlier quoted context omitted.

Oh, Anthropic, the company that hoover'd up everyone else's data, and is now unhappy when others are doing to it what it did to others? The same Anthropic?

Yes, this joke/point has been made 10,000 times in this thread in almost every comment, and on every other previous thread. Thank you!

If Anthropic doesn't like people repeating this point, Anthropic should stop repeating that they are somehow entitled to keep what they have rightfully stolen.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#437
post #408

Hypocrisy is a form of corruption. Anthropic's IP was created by harvesting and "distilling" other people's IP. Copyrighted materials, and the commons... which they have essentially privatized. The commercial goal is to avoid competition. One of the main worries for AI is "commoditization" which has come to mean "not a monopoly." To that end, it doesn't matter is the competitor is Chinese American or other. Their mot…

"Copyright violation of a published work" and "stealing private trade secrets" are in fact very different crimes. Humans have spent millenia harvesting and distilling each other's IP - "the shoulder of giants" and all that, so it's an especially disingenuous take.

> Humans have spent millenia harvesting and distilling each other's IP

You maybe somewhat correct, but also copyright lawyers wouldn’t have work if it would be up for grabs to take others IP willy nilly just because “shoulders of giants and all that”.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#438

Someone should setup a plugin or something for Claude Code that makes it easy to log all inputs and outputs for people who are willing and interested in sharing their usage. I don't want Anthropic to be the only company that can train on my usage, I want to share my usage so it can be used for training all new models. Once you have a system for collecting all logs, you just need a place where they can be submitted. I…

I don't mind being paid for using LLM, but working for free?

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#440
post #431
post #408

Hypocrisy is a form of corruption. Anthropic's IP was created by harvesting and "distilling" other people's IP. Copyrighted materials, and the commons... which they have essentially privatized. The commercial goal is to avoid competition. One of the main worries for AI is "commoditization" which has come to mean "not a monopoly." To that end, it doesn't matter is the competitor is Chinese American or other. Their mot…

> Anthropic's IP was created by harvesting and "distilling" other people's IP. Copyrighted materials, and the commons... which they have essentially privatized. Anthropic and others argue that because LLMs don’t output full copyrighted works word for word - hence their LLMs aren’t infringing on copyright laws. I think (if this ever comes to that) Chinese lab should use same arguments against Anthropic. UPDATE: this i…

> LLMs don’t output full copyrighted works word for word

Apparently they do, as per the evidence in the NYT vs OpenAI suit.

Post reply on HN