Live data from Hacker News

The Kimi K3 Moment

stephen.bochinski.dev

71–80 of 644 posts

Re: The Kimi K3 Moment

#71

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

Thanks for the models guys, sorry for your losses. Once this reality becomes mainstream and undeniable, surely the bubble pops and then what then. Future model development stops? Becomes private? Becomes a public effort?

Re: The Kimi K3 Moment

#72
Its worthwhile to have a quote from the article as some comment without reading:

"...I’ve been running Kimi K3 alongside Claude on my normal coding work, and for all practical purposes I can’t tell them apart. Same tasks, same quality of output, and near identical token counts to get there. I expected an open model to be sloppier or to grind through more tokens on the way to the same answer, and neither turned out to be true.

The prices are nowhere near each other. K3’s API runs $3 per million input tokens and $15 per million output. Claude’s top model costs $10 and $50 for the same units. The subscription side is even more lopsided..."

Re: The Kimi K3 Moment

#73
Well, there is the small issue of privacy policy: Kimi will train their models on your interactions if you use their subscriptions, and only with direct API usage (billed at API prices) they say they won't. Whether you trust that is another matter.

Those things do make a difference to some of us, even though nothing is black and white. In my case, I'll probably want to wait until other providers appear through OpenRouter and then I'll try to judge how much I trust them. But even if I don't trust them much, they don't train models anyway, so the likelihood of my data being used that way is smaller.

Re: The Kimi K3 Moment

#74
post #8

I tried Kimi K3 on a task I've done with every other model I use regularly ( https://swelljoe.com/post/i-let-every-agent-implement-its-ow... ) and found it chewed a lot longer on the problem and ate up almost the entirety of a 5 hour usage limit on their $19 plan. I only have the $20 plan from OpenAI and the same task, with a lot of the same implementation details as Kimi Code, only took a few minutes and consumed al…

With the obligatory disclaimer that I’m impressed with what open weight models can do, I have the same experience with all of them.

The benchmarks come out and say they’re as good as Opus from N months ago, then I use it for a complex task and it doesn’t work as well as Opus from N months ago did when on similar problems.

There’s a real wow factor when you get an open weights model to do amazing things, but in my experience the gap to the frontier models has always been bigger than the benchmarks would lead me to believe.

There can be a lot of value in having the cheaper open weight models for chewing through lower complexity tasks (non-programming in my primary use case) at a cheaper rate than OpenAI or other frontier API costs. Even with those I can measure bigger gaps to the frontier models than the benchmarks suggest.

If the benchmarks aren’t being directly gamed, there’s at least some selection happening where training data or model structures are being picked in ways to maximize public benchmark performance. All of the labs know there’s immense value in having good benchmarks to show for your model because most LLM consumers are picking based on lab provided benchmark charts, not running their own evals. Running your own evals is hard and expensive.

Re: The Kimi K3 Moment

#75

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

> Distillation “attacks” are not attacks. If "distillation attacks" happen, we have to conclude there is some value add in what model labs do. Regardless of how we feel about using existing human knowledge in the way they currently do, it's simply impractical to infer that everything that happens downstream of LLMs can not be an attack on some IP because of it. So both things can be true: a) People infringe on Anthro…

The output of Anthropic's models is not Anthropic's IP, as that would destroy their market, if Anthropic owned all the software it generated, and all the content. So distillation, which is just using those outputs is always going to exist.

Re: The Kimi K3 Moment

#76

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

> Distillation “attacks” are not attacks. If "distillation attacks" happen, we have to conclude there is some value add in what model labs do. Regardless of how we feel about using existing human knowledge in the way they currently do, it's simply impractical to infer that everything that happens downstream of LLMs can not be an attack on some IP because of it. So both things can be true: a) People infringe on Anthro…

Whether or not its legal to distil models, it is obviously morally permissible to do so.

Anthropic, OpenAI, etc do not deserve legal protection.

Re: The Kimi K3 Moment

#77

I never truly understood what the intended business model around LLMs was. Get them widespread through cheap pricing and then jacking it up? Being the only ones that had a viable product so to get the ability to extract as much value as you want from AI? I don't understand how a product that: - is interfaced with and is deeply linked to natural language, so everything you produce (sessions, history, etc) is in Markdo…

> I never truly understood what the intended business model around LLMs was.

A closely related question is “what do the American labs need to do in order to justify their enormous market valuations?”

It seems like the answer cannot possibly be “gradually improve model capability while figuring out how to better monetize inference.” The valuations are just way too high for that to be sufficient.

Surely the answer has to be “continually achieve large leaps in capability comparable to the first consumer releases of ChatGPT while also maintaining a significant capability lead over open models and new competitors.”

And does anyone think that’s going to happen? Even with state-level protection from competition (which incidentally would significantly harm the American economy), the large leaps in capability seem to be coming fewer and farther between.

Re: The Kimi K3 Moment

#78

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

Thanks for the models guys, sorry for your losses. Once this reality becomes mainstream and undeniable, surely the bubble pops and then what then. Future model development stops? Becomes private? Becomes a public effort?

The existing models are still going to exist. As hardware improves, there will be a day where it might cost a tenth of a penny to churn through 100M tokens a second of Opus 4.8. Established compute providers will invest in improving the models incrementally when margins drive them to look there.

Re: The Kimi K3 Moment

#79
post #37
post #29

Earlier quoted context omitted.

The visa that would correlate to this is the O-1 visa 20k O-1 visas were issued last FY which was mostly under the Trump admin, up from 19.5k the previous FY under the Biden admin

No it is H-1B visa. Right out of the university it is hard to recognize extraordinary talent. People like Sundar Pichai were not recognized as extraordinary right out of the university, he had to start at the bottom and rise up the ranks.

Melania got a EB-1 "extraordinary ability" immigrant visa

Re: The Kimi K3 Moment

#80

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

or the application layer - which will capture majority of the value.

yeah hardware companies make for nice stories or green numbers on Wall Street - but value will be captured by application layer.

look at history.

Post reply on HN