Live data from Hacker News

The Kimi K3 Moment

stephen.bochinski.dev

121–130 of 644 posts

Re: The Kimi K3 Moment

#121
post #73

Well, there is the small issue of privacy policy: Kimi will train their models on your interactions if you use their subscriptions, and only with direct API usage (billed at API prices) they say they won't. Whether you trust that is another matter. Those things do make a difference to some of us, even though nothing is black and white. In my case, I'll probably want to wait until other providers appear through OpenRo…

Indeed, the B2B / no-data-retention market is still going to provide plenty of business for American companies even if every hobbyist uses open-weight models.

Re: The Kimi K3 Moment

#122

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

>Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models So why didnt we have these LLMs in 2005?

Is this some form of rage bait? 2005 we hadn't the GPUs, we have today. There are other factors, but I think this is the big one. The mathematics of building an LLM are really old, we just hadn't the hardware to do the needed calculations.

Re: The Kimi K3 Moment

#123

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

>Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models So why didnt we have these LLMs in 2005?

We didn't have the compute required (GPUs powerful enough to parallelize forward and backward pass). This compute is what allows us to train from human knowledge or distillation.

Re: The Kimi K3 Moment

#124

Earlier quoted context omitted.

Almost all markets depend on some form of regulation whether its as simple as "leave everyone alone but no stealing" or "every participant has to source every object through mountains of red tape." Thus far the US has not really chosen to go the Chinese rare-earth method yet. The problem with distillation attacks is the end result is everyone who is not doing them is going to deal with some kind of regulation whether…

How is distillation an "attack" but gigascraping the Internet to the point of crashing servers and everyone needs Cloudflare and Anubis now not an "attack"? I'm not aiming for a what about kickflip here: I'm saying we need to either agree on some rules or stop crying foul. Maybe the coherent legal theory is that neural networks and intellectual property don't interact. That would be weird but it would be consistent,…

I think you've basically got the legal theory. Training a neural network isn't prohibited by copyright law so if you can legally get your hands on something (e.g. by sending a GET request to someone with rights to serve the contents of their web page, or by buying a book) without signing a contract to not train on it, you can train on it.

But the American AI companies only let you query their models if you first sign a contract to not train on the output.

It's hypocrisy and unfair, but I think there's a strong legal argument for it.

Of course China can simply decline to assist in enforcing that contract... But I would expect US courts to do their best to.

Re: The Kimi K3 Moment

#125

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

>Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models So why didnt we have these LLMs in 2005?

because you had neither the chips or the information in 2005. You have probably on the order of 5000x to 10000x more GPU compute today than you had in 2005 and three to four magnitudes more openly available data.

The first "L" in LLM does the work. In 2005 you had no Github, Stackoverflow, Youtube, common crawl and no archive of digital ebooks.

Re: The Kimi K3 Moment

#126

Earlier quoted context omitted.

In public with budgets that don't risk destroying the American economy presumably. Yes it may be slower.

> with budgets and what will fund these budgets exactly? inference is cheap, distillation is cheap, training is what's expensive.

Same people who fund linux kernel development. A coalition of companies that find it useful.

Re: The Kimi K3 Moment

#127

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

>Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models So why didnt we have these LLMs in 2005?

Answer the question "how much does 5 cents of LLM computation in July 2026 cost in July 2005" and you'll have the answer to your question.

Don't forget to account for all the costs. It's not just that CPUs are X times slower. Memory is X times smaller, too, and networks are X time slower. And all this hardware is many times more expensive.

If I'm getting my mental estimation right, training a 2026-frontier-class LLM in 2005 would be somewhere on the order of all the computation power in the world at the time. It's not that many more factors of magnitude before you end up at "all the computation power in the world up to that point".

Re: The Kimi K3 Moment

#128

Earlier quoted context omitted.

> western governments Are you talking about the US, specifically? Why would other countries, that don't share the same anxiety about China as the US, would be troubled with the this?

The US might pressure them?

Those days are gone. Look no further than the occupant in the White House. IE the Swedish jet industry is about to get bigger, future drone expertise, if the Ukraine can hold on if you want to learn the ins and outs, you don’t need the United States. If you’re serious about learning and building drones.

It’s going to be a different world, a world where many former allies are not gonna look to the United States first they can no longer afford to.

Re: The Kimi K3 Moment

#129
GPT 5.6 Sol comes out ahead of Kimi K3 on price/task (but not significantly so). You're probably thinking, "Why use Kimi K3? Isn't an open model supposed to beat the closed one on price?", but you need to consider that the closed models are completely hobbled when trying to do anything security-related. For my use-case, I can't risk getting pwned because I'm using a model that refuses to secure my app while there is now an open model that obliges to obliterate any app that isn't protected.

Re: The Kimi K3 Moment

#130
post #124

Earlier quoted context omitted.

How is distillation an "attack" but gigascraping the Internet to the point of crashing servers and everyone needs Cloudflare and Anubis now not an "attack"? I'm not aiming for a what about kickflip here: I'm saying we need to either agree on some rules or stop crying foul. Maybe the coherent legal theory is that neural networks and intellectual property don't interact. That would be weird but it would be consistent,…

I think you've basically got the legal theory. Training a neural network isn't prohibited by copyright law so if you can legally get your hands on something (e.g. by sending a GET request to someone with rights to serve the contents of their web page, or by buying a book) without signing a contract to not train on it, you can train on it. But the American AI companies only let you query their models if you first sign…

Contract law is never going to prevent this.
Post reply on HN