Live data from Hacker News

The Kimi K3 Moment

stephen.bochinski.dev

31–40 of 644 posts

Re: The Kimi K3 Moment

#31
post #8

I tried Kimi K3 on a task I've done with every other model I use regularly ( https://swelljoe.com/post/i-let-every-agent-implement-its-ow... ) and found it chewed a lot longer on the problem and ate up almost the entirety of a 5 hour usage limit on their $19 plan. I only have the $20 plan from OpenAI and the same task, with a lot of the same implementation details as Kimi Code, only took a few minutes and consumed al…

It would be really interesting to redo the public benchmarks for kimi k3 but token normalize the costs. Ok so maybe k3 beats fable on terminal bench, but how many tokens did it use?

Re: The Kimi K3 Moment

#32
post #22
post #8

I tried Kimi K3 on a task I've done with every other model I use regularly ( https://swelljoe.com/post/i-let-every-agent-implement-its-ow... ) and found it chewed a lot longer on the problem and ate up almost the entirety of a 5 hour usage limit on their $19 plan. I only have the $20 plan from OpenAI and the same task, with a lot of the same implementation details as Kimi Code, only took a few minutes and consumed al…

Is Kimi K3 subsidized as hard as the other models out there?

Not sure how the economics work for the Chinese models, but DeepSeek did the same task for a dime.

Re: The Kimi K3 Moment

#33
Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a cheaper version of it. There was never any plausible explanation for why this wouldn’t happen. There was never any practical mechanism to prevent someone from saving a conversation and using it to train their own model.

Even if it didn’t happen here, it was still the case that it was going to happen going forward. It was always going to end like this. Invest in the hardware companies, not the model companies.

Re: The Kimi K3 Moment

#34
post #27
post #22

Earlier quoted context omitted.

Is Kimi K3 subsidized as hard as the other models out there?

Does it matter? As an end user I really only care about 1) how much I can do in a week, and 2) how long each task takes. Subsidies would affect 1, but not 2. But if some VC wants to subsidize my Claude or Codex or whatever, awesome.

It doesn't matter if you can switch easily. It might matter if there are barriers to switching.

Re: The Kimi K3 Moment

#37
post #29
post #24

The current administration's immigration policy isn't helping. This wouldn't have happened 10 years ago because the US was this city on the hill that everyone wanted to immigrate to. Talented Asian researchers would have immigrated to the US and China would be deprived of talent.

The visa that would correlate to this is the O-1 visa 20k O-1 visas were issued last FY which was mostly under the Trump admin, up from 19.5k the previous FY under the Biden admin

No it is H-1B visa. Right out of the university it is hard to recognize extraordinary talent. People like Sundar Pichai were not recognized as extraordinary right out of the university, he had to start at the bottom and rise up the ranks.

Re: The Kimi K3 Moment

#38
post #27
post #22

Earlier quoted context omitted.

Is Kimi K3 subsidized as hard as the other models out there?

Does it matter? As an end user I really only care about 1) how much I can do in a week, and 2) how long each task takes. Subsidies would affect 1, but not 2. But if some VC wants to subsidize my Claude or Codex or whatever, awesome.

The more important question than subsidy is what is the tokenomics of running the model. If it's inefficient to run on an nvl72 cluster (or whatever the heck has enough vram to run a 3T parameter model), and k3 isn't very token efficient, then it might not be that compelling of an open weights model.

Re: The Kimi K3 Moment

#39
post #24

The current administration's immigration policy isn't helping. This wouldn't have happened 10 years ago because the US was this city on the hill that everyone wanted to immigrate to. Talented Asian researchers would have immigrated to the US and China would be deprived of talent.

- thats not a sustainable strategy - china’s homegrown tech industries already achieved escape velocity from it a long time ago, after China fenced off its market for Alibaba and Baidu in the ‘00s. some of their AI innovation at the edges was already top class 10 years ago

It has been a sustainable strategy for the tech industry for decades.

Re: The Kimi K3 Moment

#40
post #25

Earlier quoted context omitted.

https://deepswe.datacurve.ai/ or https://artificialanalysis.ai/ pareto frontier graph.

Thanks! What is the parento frontier?

The set of models that are pareto-optimal, IE for some set of variables, no other model strictly dominates them = no other model is better than them on every variable.

So like, on a cost-intelligence graph, the cheapest and most intelligent models are pareto optimal. Then in-between those if you have

- cost $3 intelligence 6

- cost $1 intelligence 5

- cost $2 intelligence 4

The 1st and 2nd are pareto optimal, the 3rd is not, because it's dominated by the 2nd (2nd is cheaper AND more intelligent at the same time)

Post reply on HN