Live data from Hacker News

Anthropic says Alibaba illicitly extracted Claude AI model capabilities

reuters.com

661–670 of 1001 posts

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#662
post #472

Earlier quoted context omitted.

Stupid question: I was under the impression that these models were trained on PB of data. Surely the amount of questions/response they can extract from querying a bigger model (Claude) is fairly modest. How is it not a drop vs the training dataset?

It's not about how big your dataset is - it's about how you use it. I jest, but I'm also completely serious. 1T tokens from Claude can teach a model something 1T tokens scraped from the open web can't. Things like "how an LLM can problem solve effectively", or "how an LLM should use tools", or "how to construct reasoning chains", or "when to double check", or "what innate capabilities an LLM can or can't rely on". Th…

Can you back up this with hard data and evidence?

Most research converges to the idea that RL on synthetic data makes models worse, not better.

If what you claim was anywhere near that relevant, than we would've long achieved singularity by simply feeding increasingly better output to the training of the next model in a loop. Yet this doesn't work.

25 million turns on Claude output is a small amount, yet an expensive one (we talking hundreds of $ millions) that is better spent on compute.

There's no evidence such a process works, but I'd like to know more if I'm wrong.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#663
post #547
post #531

Earlier quoted context omitted.

Yeah, the whole AI industry is just people ripping off each other.. Started by AI companies gulping up all the information that technical or altruistic people shared on the Internet in the past 40 years to help other fellow humans, then moved to AI companies consuming pirated and copyrighted material and now its AI companies ripping off each other. Information really does want to become free, but AI companies want to…

I'm very pro distillation. I think there needs to be distillation non profits who curate massive corpi of super high value training data from frontier models. They could have an "anonymous contribution" system where regular people with max subscriptions upload their conversation histories. It's a rough concept, but surely would be a huge boon to humanity.

sort of sounds like "project tapestry" by Yann LeCunn. Build projected data silos of highly valuable information, train in a distributed manner and share the weights upwards where they're combined and fine tuned.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#664

Earlier quoted context omitted.

If saying “plz don’t distill me” is your moat, you don’t have a moat.

No. What will happen is it will turn dark. No public release. National Security uses only, or in carefully vetted industry settings.

Good luck not crashing the markets and the economy.

And good luck not staying behind when you can't monetize your gargantuan investments and have little incentives to make your models better as the world moves on.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#665
post #465

Earlier quoted context omitted.

For something to be a trade secret, you have to actually keep it secret. If I get the ingredients of Coca-cola from an ex-employee, I've stolen a trade secret. If I work it out by doing a chemical analysis, I've stolen nothing. There is a difference with anthropic, as no-one signs a licence agreement to buy a coke. But Anthropic are also not saying you can't publish the output of their models. It's not clear to me if…

Wait, really? So why doesn't someone just reverse-engineer Coca-Cola like that? My understanding was that a "clean room" implementation is fine, but not reverse-engineering. If you can just copy everything on the market, why isn't someone already doing that?

In the case of coca cola, because use of coca leaves is highly regulated due to the fact that they also contain cocaine. There is a YouTuber who claims to have reverse engineered Coca-Cola, but he had to use tea-tree oil instead of actual coca leaf extract.

Here's EFF on reverse engineering and the law: https://www.eff.org/issues/coders/reverse-engineering-faq

Historically a lot of competition in physical products was very much reverse engineering. Because you can buy them without signing your rights away. That's why companies are keen on patents and click-through agreements.

If you look at how "clean room" processes work, they are actually a form of reverse engineering. Also clean room technique exists to avoid your new implementation infringing copyright, not trade secrets.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#667
post #540

Earlier quoted context omitted.

Yet they did not need to destroy the models which were trained with them?

Using them was allowed as fair use – it was the downloading of the pirated copies that was infringement. That's why Anthropic switched to scanning paper books.

> Using them was allowed as fair use

That is only relevant in the US, and even there it is still not clear-cut whether the fair use doctrine applies on all these scenarios. Outside of the US the situation is also quite different: for example take a look at the recent ruling on GEMA vs OpenAI in Germany.

The reality is that the copyright issue with generative AI is very complex and reaching anything resembling a conclusion will take much more than a few opinion paragraphs from an American district judge.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#668
post #637

Earlier quoted context omitted.

The compute deficit of Chinese Ai companies is real, and it IS THE ONLY competitive advantage that Western companies have. The only way the U.S. keeps that edge is to prevent distillation. The only way Chinese companies can make up for the deficit in compute is to distill. There innovation in great supply on every side of the Ocean. Its about the chips. And in terms of national security, for the U.S., and for China,…

Define compute deficit? They've been bringing out open weight models competitive with frontier models. How could they do that if they had a compute deficit?

I believe this article is about the technique they may or may not have used.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#670

Earlier quoted context omitted.

Using them was allowed as fair use – it was the downloading of the pirated copies that was infringement. That's why Anthropic switched to scanning paper books.

Isn't scanning also a form of copyright infringement? You are making a digital copy of a book, which is the same thing as downloading a book from the internet...

Copyright protects the presentation of knowledge, not the knowledge itself, which is uncopyrightable in almost all jurisdictions.

As long as the book was a legal copy, that is allowed legally.

Post reply on HN