Live data from Hacker News

The Kimi K3 Moment

stephen.bochinski.dev

191–200 of 644 posts

Re: The Kimi K3 Moment

#191
post #85

Earlier quoted context omitted.

Given these models could not have been trained in the first place if they had to license every line of random fan fiction on the internet, I think distillation also being fair game is a tradeoff everyone should be willing to take (unless they want to decelerate, but that's a different conversation).

Us models didnt pay for licenses too

We're still in the early days of the AI industry timeline(relative to traditional industries). Not everything has yet been litigated.

Taxes on AI subscriptions or AI capable hardware, to financially compensate IP holders for (potential) IP theft, could very well arrive in the near future, once the industry is mature.

If this shocks you and sounds preposterous, I'll remind you that in several EU countries, we still pay extra taxes on any and all storage mediums and on devices with built-in storage (tapes, CDs, DVDs, HDDs, SSDs, tablets, phones, etc) simply because they can be used to store pirated content, decisions based on laws from 50-100 years ago, and the money goes to the national unions and associations of music and arts IP holders. It's basically a lobby pushed and government legalized extortion racket that no voter agrees with or can change but has no choice but to conform either way.

So I guarantee you in the future, it will be the same for AI subscriptions and hardware capable of running LLMs locally. Every time you purchase a Claude or ChatGPT subscription, an Nvidia GPU, Intel/AMD SoC PC or an Apple/Qualcomm powered smartphone, you'll pay a government enforced tax to the likes of Sony, Axel Springer, etc. for licensing their IP, whether you want to or not. In the EU at least. US maybe not.

Re: The Kimi K3 Moment

#192
post #85

Earlier quoted context omitted.

Given these models could not have been trained in the first place if they had to license every line of random fan fiction on the internet, I think distillation also being fair game is a tradeoff everyone should be willing to take (unless they want to decelerate, but that's a different conversation).

Us models didnt pay for licenses too

That is incorrect. Anthropic paid $1.5 billion in compensation to copyright holders for use of their content in training data. OpenAI pays hundreds of millions per year across 150+ licensing deals for access to copyrighted data. Meta and Alphabet have similar arrangements.

Under the settlement, Anthropic was forced to delete the pirated data they were training on.

Chinese labs can still train on pirated data. I doubt the Chinese models operate under similar licensing agreements.

Re: The Kimi K3 Moment

#193
post #112

Earlier quoted context omitted.

I was more thinking they would be funding US labs.

the question was: what is the endgame for the stated "second class labs" strategy of distilling their frontier competitors then undercutting them on price?

Yes yes, we all understand the game-theoretic race-to-the-bottom you're describing here. Somehow despite linux being FOSS it still powers most of the important computing in the world. Can you explain how that works despite it being free? Once you understand that case I think you'll understand the game-theory behind how large projects can exist in the absence of traditional IP protection.

Re: The Kimi K3 Moment

#194

Earlier quoted context omitted.

Almost all markets depend on some form of regulation whether its as simple as "leave everyone alone but no stealing" or "every participant has to source every object through mountains of red tape." Thus far the US has not really chosen to go the Chinese rare-earth method yet. The problem with distillation attacks is the end result is everyone who is not doing them is going to deal with some kind of regulation whether…

How is distillation an "attack" but gigascraping the Internet to the point of crashing servers and everyone needs Cloudflare and Anubis now not an "attack"? I'm not aiming for a what about kickflip here: I'm saying we need to either agree on some rules or stop crying foul. Maybe the coherent legal theory is that neural networks and intellectual property don't interact. That would be weird but it would be consistent,…

Knowlege should not have ownership. Training and distillation should be allowed

Granting people some form of control over knowledge only serves the public interest inasmuch it provides incentive to create more of it. Mass media, effortless duplication, and copyright extensions had already broken this to the point where control of knowledge was suppressing creation of new knowledge more than it facilitated.

The world has changed, we need a mechanism that works for the public interest that applies to the facts as they now are.

Re: The Kimi K3 Moment

#195

Earlier quoted context omitted.

> Distillation “attacks” are not attacks. If "distillation attacks" happen, we have to conclude there is some value add in what model labs do. Regardless of how we feel about using existing human knowledge in the way they currently do, it's simply impractical to infer that everything that happens downstream of LLMs can not be an attack on some IP because of it. So both things can be true: a) People infringe on Anthro…

Regardless of whether it’s intellectual property or it isn’t intellectual property, it doesn’t actually matter. If AI doesn’t stop seeing diminishing returns in scaling up, and it hasn’t yet in the 10 years since the attention/transformers paper, the advent of AI will be the most important development in the history of humanity. Controlling that machine, or at least having one of your own, is an existential problem f…

It's unlikely the USA would be granted an exclusive patent for the atomic bomb given the well-established existence of prior-art in the form of nuclear fission on the sun.

(I actually appreciated your analogy, despite my lark)

Re: The Kimi K3 Moment

#196
post #124

Earlier quoted context omitted.

How is distillation an "attack" but gigascraping the Internet to the point of crashing servers and everyone needs Cloudflare and Anubis now not an "attack"? I'm not aiming for a what about kickflip here: I'm saying we need to either agree on some rules or stop crying foul. Maybe the coherent legal theory is that neural networks and intellectual property don't interact. That would be weird but it would be consistent,…

I think you've basically got the legal theory. Training a neural network isn't prohibited by copyright law so if you can legally get your hands on something (e.g. by sending a GET request to someone with rights to serve the contents of their web page, or by buying a book) without signing a contract to not train on it, you can train on it. But the American AI companies only let you query their models if you first sign…

> It's hypocrisy and unfair, but I think there's a strong legal argument for it.

That right there is the problem.

Re: The Kimi K3 Moment

#197
post #136

Earlier quoted context omitted.

Look how hard Anthropic is to even be able scroll back on your conversation, or look at the thinking tokens or subagents. They want to keep everyone coming back to the watering hole but never to learn how to dig a well.

Why is it hard to scroll?

It really is not, not sure what OP is on about.

Re: The Kimi K3 Moment

#198

Earlier quoted context omitted.

> Distillation “attacks” are not attacks. If "distillation attacks" happen, we have to conclude there is some value add in what model labs do. Regardless of how we feel about using existing human knowledge in the way they currently do, it's simply impractical to infer that everything that happens downstream of LLMs can not be an attack on some IP because of it. So both things can be true: a) People infringe on Anthro…

>People infringe on Anthropics IP No. Authors do not infringe on IP when they read another's book, nor should the lumber company be able to dictate how I use planks and if I can resell them if i'm done with them. You're framing it as if the added value of the author or lumber company, awards them consideration when somebody uses the products to create more value. IP law was always a big mess, and these questions cros…

It's more simple: They infringe on the IP by way of violating the ToS. If you violate ToS and the company suffers financial harm, they usually can (usually) sue you in civil court for damages.

Re: The Kimi K3 Moment

#199

Earlier quoted context omitted.

Us models didnt pay for licenses too

That is incorrect. Anthropic paid $1.5 billion in compensation to copyright holders for use of their content in training data. OpenAI pays hundreds of millions per year across 150+ licensing deals for access to copyrighted data. Meta and Alphabet have similar arrangements. Under the settlement, Anthropic was forced to delete the pirated data they were training on. Chinese labs can still train on pirated data. I doubt…

they didn't pay yet, because court challenged settlement as inadequate.

> I doubt the Chinese models operate under similar licensing agreements.

US corps likely pay licenses when afraid to be sued, or have troubles getting that data, otherwise they just take data, which was demonstrated many times. The same apply to Chinese corps, alibaba totally can be sued in US.

Re: The Kimi K3 Moment

#200

Earlier quoted context omitted.

Us models didnt pay for licenses too

That is incorrect. Anthropic paid $1.5 billion in compensation to copyright holders for use of their content in training data. OpenAI pays hundreds of millions per year across 150+ licensing deals for access to copyrighted data. Meta and Alphabet have similar arrangements. Under the settlement, Anthropic was forced to delete the pirated data they were training on. Chinese labs can still train on pirated data. I doubt…

After the fact. They did the same thing Youtube, Uber and Airbnb did: Break the law, eventually get caught, cut some deal where they pay a pittance and keep doing the same thing but now with leverage on their side.
Post reply on HN