Live data from Hacker News

Anthropic says Alibaba illicitly extracted Claude AI model capabilities

reuters.com

671–680 of 1001 posts

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#671

Earlier quoted context omitted.

The AI companies seem to take the viewpoint that everything on the internet is free, except their stuff. It's okay to hammer some random website with AI crawlers, ignoring robots.txt, and causing bandwidth costs to skyrocket. But if you cost an AI provider money with your data acquisition practices, well, that's just clearly unacceptable.

> But if you cost an AI provider money with your data acquisition practices, well, that's just clearly unacceptable. It's the same question libertarian advocates cannot resolve: If one truly believes in personal sovereignty, how are shared resources paid for, such as roads, power grids, potable water, sewage services, fire departments, and police departments? It is also not a coincidence that leadership in many tech…

Libertarians can just flip it round and say how do socialists solve the free rider problem? Neither system resolves both problems.

Extremist dogma is not a great way to run a society, but it does good numbers on social media, so here we are.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#672
post #534

I'll just leave it here: "Anthropic's downloading of over seven million books from pirate sites like LibGen constituted infringement, the judge ruled, rejecting Anthropic's "research purpose" defense: "You can't just bless yourself by saying I have a research purpose and, therefore, go and take any textbook you want." https://www.joneswalker.com/en/insights/blogs/ai-law-blog/wh...

In the early days of music streaming, many of the entrants were seeding their service with vast libraries of pirated content. The winners cut deals with the copyright holders and then went after the rest.

Or the early days of video uploads, YouTube's most watched videos were "pirated" clips from popular shows (e.g. SpongeBob, The Daily Show) and part of the reason I went to YouTube instead of other video hosting sites (e.g. DailyMotion).

Viacom sued YouTube, while CBS and Universal ended up licensing their content.

https://www.eff.org/deeplinks/2007/03/viacom-v-google-invest...

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#674

The hypocrisy of Anthropic complaining about "illicitly extracting its Claude AI model capabilities" and supporting the White House's accusation of China "stealing U.S. AI labs' intellectual property on an industrial scale" is hilarious. Anthropic, OpenAI, Google, Microsoft, et al trained their models by ignoring the rights of copyright holders when harvesting whatever content they could. Now one of them is crying fo…

Not really.

Data mining for AI is presumably fair use, whereas when you sign up for a Claude account, you enter into a legally binding contract that says you will not distill a model based on its outputs.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#675
post #534

I'll just leave it here: "Anthropic's downloading of over seven million books from pirate sites like LibGen constituted infringement, the judge ruled, rejecting Anthropic's "research purpose" defense: "You can't just bless yourself by saying I have a research purpose and, therefore, go and take any textbook you want." https://www.joneswalker.com/en/insights/blogs/ai-law-blog/wh...

In the early days of music streaming, many of the entrants were seeding their service with vast libraries of pirated content. The winners cut deals with the copyright holders and then went after the rest.

[deleted]

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#676
Seems like a fair play by Alibaba. However, is there any "open source" attempt at crowdsourcing distillation?

Like some place people can submit their chatbot convos so they can be aggregated?

Like an equivalent to OpenCrawl but for mining the models. It feels like thatd be a richer dataset than Alibaba generating queries and feeding them into Anthropic/OpenAI models

PS: Does anyone know how when companies distill each others' models the synthetic queries are generated? Im just assuming theyd be worse than organic ones

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#677

Earlier quoted context omitted.

It's not about how big your dataset is - it's about how you use it. I jest, but I'm also completely serious. 1T tokens from Claude can teach a model something 1T tokens scraped from the open web can't. Things like "how an LLM can problem solve effectively", or "how an LLM should use tools", or "how to construct reasoning chains", or "when to double check", or "what innate capabilities an LLM can or can't rely on". Th…

Unremarkable base model will remain an unremarkable fine-tuned model that memorised a couple thousand of input-output pairings.

Ha ha, as if.

Base models have a lot of capabilities - arranged in all the wrong ways for high performance reasoning and problem-solving. The power of fine tuning on "a couple thousand of input-output pairings" is that it can fix some of that. If your pairings are very well chosen, that is.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#679
post #672

Earlier quoted context omitted.

In the early days of music streaming, many of the entrants were seeding their service with vast libraries of pirated content. The winners cut deals with the copyright holders and then went after the rest.

Or the early days of video uploads, YouTube's most watched videos were "pirated" clips from popular shows (e.g. SpongeBob, The Daily Show) and part of the reason I went to YouTube instead of other video hosting sites (e.g. DailyMotion). Viacom sued YouTube, while CBS and Universal ended up licensing their content. https://www.eff.org/deeplinks/2007/03/viacom-v-google-invest...

They still are. My kids haven't watched a single Simpsons or Family Guy episode but are quoting both regularly.

Facebook et al also quite literally stole email contact lists and installed spyware at kernel level on mobile phones which they used to spy on all Android users. Via the phone manufacturers.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#680

Here's what is happening: Chinese resellers are offering Claude tokens at 70-90% below official Anthropic API prices. They achieve this by reselling capacity from pooled Claude Max accounts, payments fraud, and also reselling the model output & reasoning chains to various Chinese labs. They are subsidizing model access in exchange for user logs and reasoning traces, which they then sell as training data, allowing the…

Why aren't these on openrouter?

Probably because Openrouter is a US based company, and they don't want to be sued by Anthropic/OpenAI.
Post reply on HN