Live data from Hacker News

Kimi K2.6: Advancing open-source coding

kimi.com

271–280 of 394 posts

Re: Kimi K2.6: Advancing open-source coding

#271
Early benchmarks show tremendous improvement over Kimi K2 Thinking, which didn't perform well on our benchmarks (and we do use best available quantization).

Kimi K2.6 is currently the top open weights model in one-shot coding reasoning, a little better than GLM 5.1, and still a strong contender against SOTA models from ~3 months ago (comparable to Gemini 3.1 Pro Preview).

Agentic tests are still running, check back tomorrow. Open weights models typically struggle with longer contexts in agentic workflows, but GLM 5.1 still handled them very well, so I'm curious how Kimi ends up. Both the old Kimi and the new model are on the slower side, so that's a consideration that makes them probably less usable for agentic coding work, regardless. The old Kimi K2 model was severely benchmaxxed, and was only really interesting in the context of generating more variation and temperature, not for solving hard problems. The new one is a much stronger generalist.

Overall, the field of open weights models is looking fantastic. A new near-frontier release every week, it seems.

Comprehensive, difficult to game benchmarks at https://gertlabs.com/?mode=oneshot_coding

Re: Kimi K2.6: Advancing open-source coding

#272

Earlier quoted context omitted.

It's been released to "select partners".

Yeah, Crowdstrike among them. Clearly experts in this "security" thing, given what happened during the last incident...

Yeah, people who would look stupid if they said the king had no clothes.

Re: Kimi K2.6: Advancing open-source coding

#273
post #7

Earlier quoted context omitted.

With the previous generation? Yes. With 10T mythos-level models? Not even close.

Mythos isn't the current generation, it's literally vaporware.

I doubt it's literal vaporware. It's likely just a variant of whatever model they just generally released with some fancy prompt and a highr quant.

T

Re: Kimi K2.6: Advancing open-source coding

#274

Early benchmarks show tremendous improvement over Kimi K2 Thinking, which didn't perform well on our benchmarks (and we do use best available quantization). Kimi K2.6 is currently the top open weights model in one-shot coding reasoning, a little better than GLM 5.1, and still a strong contender against SOTA models from ~3 months ago (comparable to Gemini 3.1 Pro Preview). Agentic tests are still running, check back t…

Surprised to see such variance per language

Re: Kimi K2.6: Advancing open-source coding

#275

Has anyone here used Kimi for actual work? I tried it once, although it looks amazing on benchmarks, my experience was just okay-ish. On the other hand, Qwen 3.6 is really good. It’s still not close to Opus, but it’s easily on par with Sonnet.

Before GLM-5.1, I was going back and forth between Opus 4.5 and Kimi 4.5 and having very good results with Kimi.

Re: Kimi K2.6: Advancing open-source coding

#277

Earlier quoted context omitted.

At this point drawing these Pelicans must be in the training data sets.

not if I can help it! https://github.com/scosman/pelicans_riding_bicycles

That's truly a wonderful collection of pelicans riding bicycles.

Much Win! ;)

Re: Kimi K2.6: Advancing open-source coding

#278

There is some humor in the fact that china (of all countries) is pioneering possibly the world's most important tech via open source, while we (US) are doing the exact opposite.

A lot of people speculating on the motivations behind Chinese labs open sourcing their models. The reason is simple and clear: It is the only viable commercialization strategy that is available to them. I wrote about this here: https://try.works/writing-1#why-chinese-ai-labs-went-open-an...

Re: Kimi K2.6: Advancing open-source coding

#279
post #208

Earlier quoted context omitted.

First they came for people asking about Tiananmen Square And I did not speak out Because I was not asking about Tiananmen Square Then they came for people asking about Israel And I did not speak out Because I was not asking about Israel

This made me chuckle. I didn't mean to dismiss ethical accountability for LLM training corpuses. It is a shame. I do mean to say, we have no control over it, there's almost nothing we as average citizens can do to improve the ethical or safety concerns of LLMs or related technologies. Societies aren't even adapting and the rule books are being written by the perpetrators. Might as well get out of it what we can while…

Wonder if stuff like this would affect it?

https://github.com/p-e-w/heretic

Guessing it probably would?

Re: Kimi K2.6: Advancing open-source coding

#280
post #129

There is some humor in the fact that china (of all countries) is pioneering possibly the world's most important tech via open source, while we (US) are doing the exact opposite.

I wonder if there's a strategy behind all of this on China's side. I know the CCP uses a direct hand in many affairs in China, but is there an actual coordinated effort to compete with, or sabotage the West?

Chinese labs have no marketing and sales capacity in the overseas market, so they in fact have no choice but to open source their models as that is what brings awareness and trust in their models. In fact, it is overseas open source marketing that drives adoption of their models in China as well. I wrote about this here: https://try.works/writing-1#why-chinese-ai-labs-went-open-an...
Post reply on HN