Live data from Hacker News

Anthropic says Alibaba illicitly extracted Claude AI model capabilities

reuters.com

421–430 of 1001 posts

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#422

This is great for competition! Chinese vendors offering a cheaper solution = what economics told me the free market was all about. I also learnt that Anthropic should get better at what they do if they want to compete. If not, somebody else will win. Or does this not apply to huge US corporations any more?

The "free market" gave the PRC its current strategic lock on rare-earth minerals. There's definitely no such thing as a free market in a Maoist dictatorship. I personally think the "free market" concept is an unachievable ideal and thought-terminating cliche, but "free market in a Maoist dictatorship" is for sure a contradiction in terms.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#425
post #384

Someone should setup a plugin or something for Claude Code that makes it easy to log all inputs and outputs for people who are willing and interested in sharing their usage. I don't want Anthropic to be the only company that can train on my usage, I want to share my usage so it can be used for training all new models. Once you have a system for collecting all logs, you just need a place where they can be submitted. I…

Yikes, no thank you

Do you have a substantive reason why you dislike this? What is the problem if it's opt-in? Nobody is forcing you to share your usage if you don't want to.

I'd prefer it if all the model builders could train on my usage rather than being limited to a single company. That'll hopefully help make all the models better in the long-term.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#426
post #408

Hypocrisy is a form of corruption. Anthropic's IP was created by harvesting and "distilling" other people's IP. Copyrighted materials, and the commons... which they have essentially privatized. The commercial goal is to avoid competition. One of the main worries for AI is "commoditization" which has come to mean "not a monopoly." To that end, it doesn't matter is the competitor is Chinese American or other. Their mot…

"Copyright violation of a published work" and "stealing private trade secrets" are in fact very different crimes.

Humans have spent millenia harvesting and distilling each other's IP - "the shoulder of giants" and all that, so it's an especially disingenuous take.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#427

Earlier quoted context omitted.

But I can rebuild glm Using open source methods…

And there are a ton of Claude conversation logs (with CoT/inference) with no clear provenance circulating freely on huggingface, guess where they (likely) come from.

Wouldn't all of them be in Mandarin according to this theory? Are they?

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#428
post #140

Earlier quoted context omitted.

>This is great for competition! Chinese vendors offering a cheaper solution = what economics told me the free market was all about. Yeah, like all those Chinese bootleggers selling DVDs for a few dollars rather than $20. Free market! https://news.ycombinator.com/item?id=48664814

"Information wants to be free" Anthropic profited from training its models on all kinds of copyrighted information, live by the sword, die by the sword... Their model weights, training data, training methods, etc are all going to leak to China over time. Nobody on a site named _Hacker_ news should be all that upset about this.

Don't forget insider threat vector, too.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#430
post #213

There's two basic kinds of distillation: 1) the massive [and dumb] method where you ask a question and use the answer as reinforcement (Black Box), and 2) more targeted distillation where you use one model to directly inform/train/guide another model (RLAIF). The latter is basically fine-tuning the model with direction from another model. Thousands of businesses do this every day to fine-tune. This is almost certainl…

If you’re doing evals, you’re basically doing RLAIF without training a model; just looking at the results. Fundamentally it is very difficult to stop this while still making your AI models useful.

Similarly, if you did a corpus study on bioRvix to summarize recent science findings — you could use the same questions and answers to fine tune a model.

There is no way to communicate information at scale to companies through the API, for anything approaching a real application, without that information forming a corpus another model can be trained on.

But it wouldn’t be the first time they broke a model:

Their “guardrails” that cause it to reject user prompts also means it relies on its pop science summary of medicine to tell you why bioRxiv is wrong rather than accurately summarize the papers.

They’ve successfully created a smug, argumentative average of the internet which refuses to even consider it might be wrong or that it’s reading a science paper which is based on measurements and not vibes — but why would I pay for that?

I get it for free online.

Post reply on HN