Live data from Hacker News

The Kimi K3 Moment

stephen.bochinski.dev

541–550 of 644 posts

Re: The Kimi K3 Moment

#541

Earlier quoted context omitted.

>through proxies and heavily discounted token resellers Could you explain a little more about how this works? Are you saying that the Chinese run or have backdoored something like OpenRouter?

Have a look at https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens... and https://x.com/yan5xu/status/2029743983522631698 Chinese resellers acquire hundreds of Claude Max 5x accounts and set up a custom proxy server. Customers point their ANTHROPIC_API_KEY at that proxy, and requests are routed to Anthropic through one of those hundreds of accounts. Because one $200 Claude Max 5x account gets the equivalent…

oh wow, an entire seedy underbelly I was unaware of. Thanks, great reply. Appreciated!

Re: The Kimi K3 Moment

#543

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

I strongly agree with the premise that distillation is not an “attack”. But that said: K3 is not a distilled version of Fable or Sol. Fable has been barely available and Sol was just released! Moreover, K3 is superior to both models in some domains, according to user scoring on the Arena. API distillation can’t give you these results anyway. All it is useful for is bootstrapping RL in new domains to get past the “col…

If your business leaks to your customers… if you can ask your product for its inner secret sauce and it readily gives away the goose. That is not a moat.

Re: The Kimi K3 Moment

#544

Earlier quoted context omitted.

You skipped the part where Hyundai chose to sell the cars at loss in the hopes of eventually gaining a monopoly position.

Misleading on both counts: 1) Anthropic tokens via subscription aren't sold at a loss, they're sold at cost. 2) Subscription plans are not sold in hopes of eventually gaining a monopoly position. They act as a loss leader designed to get a foot-in-the-door and funnel companies into costly enterprise plans, where Anthropic can charge full API rates.

[dead]

Re: The Kimi K3 Moment

#545

Earlier quoted context omitted.

I strongly agree with the premise that distillation is not an “attack”. But that said: K3 is not a distilled version of Fable or Sol. Fable has been barely available and Sol was just released! Moreover, K3 is superior to both models in some domains, according to user scoring on the Arena. API distillation can’t give you these results anyway. All it is useful for is bootstrapping RL in new domains to get past the “col…

API distillation doesn't have to explain all of K3's capabilities for it to have happened. Kimi K3 reproducibly identifies itself as Claude: https://x.com/denisewu/status/2077984660211269870 This behavior is exactly what you'd expect from a model distilled from Claude. There's a detailed analysis of K3's ambiguous identity here: https://github.com/rgreenblatt/which_claude_is_k3/blob/main/... This analysis observed K3…

Surprising they didn't clean that from the data before training. It's easy to identify, a simple search->replace gets most of it, and a cheap LLM can identify the edge cases (e.g. avoiding "Claude Shannon" -> "Kimi Shannon" or something).

Re: The Kimi K3 Moment

#546

Earlier quoted context omitted.

There are huge evidence of copying. Some day China can pioneer in science or technology but the current claim about Chinese companies leading AI development is ridiculous given the evidence of distillation and the fact that like 95 percent of science that lead to the current state of AI happened in either North America or Europe. To be honest if you want to list academic papers that lead to the current AI models the…

From the Hoover Institution’s analysis of the team behind DeepSeek: “We find striking evidence that China has developed a robust pipeline of homegrown talent. Nearly all of the researchers behind DeepSeek’s five papers were educated or trained in China. More than half of them never left China for schooling or work, demonstrating the country’s growing capacity to develop world-class AI talent through an entirely domes…

Exactly. China is a real tech power now, just like Japan and Taiwan. The U.S. is ahead in a lot of areas of technology, but China has home grown talent that is taking the lead in other areas. And unlike Japan and Taiwan, China has a much bigger pool to draw from.

Re: The Kimi K3 Moment

#547

According to OpenAI's "head of strategic futures": 1) Kimi 3 is a "very good model" 2) It's performance can NOT be explained by distillation 3) The US government should create FUD to stop US corporations from using it (so they use OpenAI instead) https://x.com/deanwball/status/2078133895766114412

Fascinating take from OpenAI. It really gives the lie to the idea that they see AI leading to a better life for all. "One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a 'public good' which will ultimately be provided by the state as a kind of 'digital public infrastructure.' This future strikes me as a dystopi…

If the government provided funding to independent organizations that made public-goods AIs, that could be great, as long as the government had no editorial control.

It sounds like he's imagining AIs only being trained and provided by governments, which could get pretty dystopian.

Re: The Kimi K3 Moment

#548
post #460

Earlier quoted context omitted.

I mean, AWS Bedrock (with the exception of Fable) gives enterprises the same assurances (but again, with the exception of Fable, which is explicitly listed as requiring data egress [or exfil, depending on how you look at it] outside of your contractual AWS security boundary).

It gives you "assurances" that can be broken. They could still be clean-rooming your prompts like Anthropic does. The only way to be certain is if they don't have access to your data to begin with.

What do you mean by "clean-rooming your prompts?"

Re: The Kimi K3 Moment

#549
post #494

Earlier quoted context omitted.

even 8x rtx pro 6000 is only 768GB of VRAM. IDK how anyone is going to run k3

on the 1.5 tB macStudio that is going to be released next quarter

768GB next quarter. The chip supporting 1.5TB won't come out until 2028 or 2029.

Re: The Kimi K3 Moment

#550
post #332

Even in this very thread the feedback on Kimi's actual efficacy is debated. I personally feel its worse than both Fable and 5.6 Sol, but I feel like the conversation isn't really about whether its good or not, but a backlash against the U.S governments foray into regulation. So I think people _want_ it to be superior out of anger/frustration with the current situation.

Just for the sake of argument and using some admittedly insane numbers, give me Opus 4.5 at a tenth the cost and running ten times as fast and I'd take that for almost any coding task over any current frontier model. There was a real phase transition somewhere in that range and improvements since then, while impressive and useful and by the benchmarks quite large, have in practice not been anywhere near as big a phas…

> give me Opus 4.5 at a tenth the cost and running ten times as fast and I'd take that for almost any coding task over any current frontier model

This is pretty much where I'm at

Post reply on HN