Live data from Hacker News

Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

thinkpol.ca

11–20 of 235 posts

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#11

In a single challenge, measured by how performant the solution was. Kimi K2.6 is definitely a frontier-sized model, so on the one hand it's not that surprising it's up there with the closed frontier models. Being open is nice though, even though it doesn't matter that much for folks like me with a single consumer GPU.

[flagged]

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#12

In a single challenge, measured by how performant the solution was. Kimi K2.6 is definitely a frontier-sized model, so on the one hand it's not that surprising it's up there with the closed frontier models. Being open is nice though, even though it doesn't matter that much for folks like me with a single consumer GPU.

It absolutely does matter.

The enshittification will go unnoticed at first but I'm already finding my favourite frontier models severely nerfed, doing incredibly dumb stuff they weren't in the past.

We need open weight models to have a stable "platform" when we rely on them, which we do more and more.

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#13
post #5
post #4

I’m a little confused as to the setup. It was asking each model to one-shot a script and then the scripts faced off? Were the models given a computer environment? Or a test server to iterate against?

Sounds incredibly simple to me. One-shot.

So nothing like real-world coding, where you’d be able to run and test the script before submitting?

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#14
What's the GPU VRAM requirements for this thing?

Awesome to have a open model that can compete, but damn it would be so much better if you could run it locally. Otherwise, it's almost so difficult to run (e.g. self host) that it's just way more convenient to pay OpenAI, Claude, etc

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#15

Great to know, but what was the cost both in terms of $$ and tokens used? Not to invalidate these benchmark results because they are useful, but the real usefulness it what they are capable to do when real people interact with them at scale. Regardless, these are good news, because now that Microsoft is basically giving up their all-in strategy with Github's Copilot and Anthropic is playing the "I'm too good for you"…

Re pricing. Never as high as frontier commercial models.

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#17

In a single challenge, measured by how performant the solution was. Kimi K2.6 is definitely a frontier-sized model, so on the one hand it's not that surprising it's up there with the closed frontier models. Being open is nice though, even though it doesn't matter that much for folks like me with a single consumer GPU.

>Being open is nice though, even though it doesn't matter that much for folks like me with a single consumer GPU.

Of course it matters because that makes coding plans much cheaper than those from Anthropic and OpenAI.

For personal use I have coding plans with GLM 5.1, Kimi K2.6, MiniMax M2.7 and Xiaomi MiMo V2.5 Pro and I am getting a lot of bang for the buck.

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#18
post #7

I have to try Kimi. I was looking for an alternative. If you have any experience, advice, please share. I saw Kimi is at the top of the Open Router ranking.

I use Kimi at home via a kimi.com subscription and Kimi CLI (sometimes running inside Zed, sometimes not). My favorite model by far. And it's just $20.

I have to use a supposedly frontier model at work and I hate it.

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#19
post #6

In a single challenge, measured by how performant the solution was. Kimi K2.6 is definitely a frontier-sized model, so on the one hand it's not that surprising it's up there with the closed frontier models. Being open is nice though, even though it doesn't matter that much for folks like me with a single consumer GPU.

This is the future though. Open weights models that run on H200s provide far more opportunity to build products and real infrastructure around. You can always distill this for your little RTX at home. But models shaped for consumer hardware will never win wide adoption or remain competitive with frontier labs. This is something that _can_ compete. And it will both necessitate and inspire a new generation of open clou…

I don’t fully understand what open weights unlocks that cannot be accomplished via API from a product standpoint.

Open weights is great if you want to do additional training, or if you need on-prem for security.

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#20
post #7

I have to try Kimi. I was looking for an alternative. If you have any experience, advice, please share. I saw Kimi is at the top of the Open Router ranking.

Kimi K2.6 is great but I advice you to get a coding plan from Kimi.com as that way is much cheaper than paying for API calls using OpenRouter.
Post reply on HN