Live data from Hacker News

Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

thinkpol.ca

191–200 of 235 posts

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#191

Earlier quoted context omitted.

Except that if you tried one-shotting your ticket twenty times at different hours of the day and different days of the week, you would have enough changes to make benchmarks even if you used the same model every time. Much moreso if you fiddled with the thinking or changed the prompt. Because non-deterministic, because of constant updates and changes, and because the models are throttled according to number of users,…

You never get "the same" Steph Curry, he might be tired, annoyed by a fan, getting older... but if he and I were to throw 100 3-pointers, we could all correctly guess who will perform better.

Good point.

But I use Codex and Claude daily (work and hobby respectively). And there are days where one or the other just seems to have gotten up on the wrong side of the bed. Or is just being lazy. Or is suddenly super-powered do everything including what i asked it not to. (To be fair, the same thing happens with myself. :/)

I am convinced that if I was bench-marking, I would be convinced these are different models on different days.

[This conviction may say more about me then about the model.]

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#192
post #179

Earlier quoted context omitted.

That is the very reason the open source models exist. Prestige and soft power to influence interest away from American models and hopefully slow down their progress.

DeepSeek and other Chinese model makers are massively accelerating progress in AI not slowing it down. They're the only ones who still come up with real technical innovations while the proprietary model makers are stagnating.

That is a petty big assumption (aka bullshit) unless you have direct insight the inner workings of the big US labs. Just because it isn’t published doesn’t mean that innovation is not happening.

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#193

I was surprised by the ranking, until I read what the test was. Not horribly relevant for coding. The current ranking of all tests makes more sense (well, except for how well Gemini does) https://aicc.rayonnant.ai

All those models and the site is not responsive on mobile. Ironic.

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#194

Earlier quoted context omitted.

DeepSeek and other Chinese model makers are massively accelerating progress in AI not slowing it down. They're the only ones who still come up with real technical innovations while the proprietary model makers are stagnating.

That is a petty big assumption (aka bullshit) unless you have direct insight the inner workings of the big US labs. Just because it isn’t published doesn’t mean that innovation is not happening.

That's an unfalsifiable assertion with no evidence to support it, while all the visible evidence we do have points to stagnation and merely incremental pushes among the big proprietary model makers. Even Claude Mythos, which was 'teased' to the public but not released, is reportedly mostly a scaled-up model that takes massive compute resources to run (and lengthy agentic loops to achieve its reported results in computer security). The polar opposite to what the Chinese labs are releasing now.

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#195

Earlier quoted context omitted.

That is a petty big assumption (aka bullshit) unless you have direct insight the inner workings of the big US labs. Just because it isn’t published doesn’t mean that innovation is not happening.

That's an unfalsifiable assertion with no evidence to support it, while all the visible evidence we do have points to stagnation and merely incremental pushes among the big proprietary model makers. Even Claude Mythos, which was 'teased' to the public but not released, is reportedly mostly a scaled-up model that takes massive compute resources to run (and lengthy agentic loops to achieve its reported results in compu…

So no insight and just going off blogposts and YouTube huh. Pot kettle calling each other black etc.

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#196

These posts are going to be a constant for the next year, because there's no objective way to compare models (past low-level numbers like token generation speed, average reasoning token amount, # of parameters, active experts, etc). They're all quite different in a lot of ways, they're used for many different things by different people, and they're not deterministic. So you're constantly gonna see benchmarks and test…

I agree. I have rather constrained use cases for LLMs and the agentic harnesses that I use with them.

I try one or two of my use cases with new models or harnesses, make my own often subjective judgements, and largely ignore benchmarks.

Blogging and writing in general are a business, or feed other tech adjacent businesses, and a lot of writing about evals is attention getting - nothing wrong with that but there is a lot of noise.

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#197
post #167

Earlier quoted context omitted.

I really don't mean to discourage anyone but I find that making your own harness is pretty complex process. We where very naive when we started, thinking this is just a loop, but then it turned out it has all kinds of nuances, edge-cases, things to be considered, etc. As I am writing this I am fixing a bug in the harness that could cause infinite loops under some conditions. For simple interactions looping over the t…

> I really don't mean to discourage anyone but I find that making your own harness is pretty complex process. We where very naive when we started, thinking this is just a loop, but then it turned out it has all kinds of nuances, edge-cases, things to be considered, etc. I'd be interested in your harness if its open source, please share some more resources > Pi.dev, Hermes and all the other I have seen do not do even…

https://github.com/can1357/oh-my-pi

Pi+some additions here. Have not used it enough to have a strong opinion, however, I’ve seen it mentioned here.

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#198

Earlier quoted context omitted.

My theory is we will end up in a similar spot to hiring people. You can look at a CV (benchmarks) but you won't know for sure until you've worked with them for six months. We as an industry cannot determine if one software engineer is objectively better than another, on practically any dimension, so why do we think we can come to an objective ranking of models?

Not many things are as manifold broken as hiring these days. I hope we do not end up there.

[deleted]

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#199
post #188
post #103

Earlier quoted context omitted.

If you look at the ranking breakdown though, Kimi K2.6 has only participated in the last 5 challenges (claude dominated before then) and if you only count those it would be in first place

It also has a DNF. So it has a high ceiling but also unfortunately a low floor. So using Kimi means accepting high variability of the output. Personally what I've found that has made coding agents more and more useful over the last year is that they have gotten a higher and higher floor, not that they have gotten a higher and higher ceiling. They were already plenty smart a year ago, it was just that they failed so o…

If you look at the last 5 challenges (the ones Kimi was in) both Claude and Kimi have 1 DNF, chatgpt has 2

I'm not sure this is enough data to form an opinion, but going by what we have Kimi would be as reliable as Claude

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#200

At the current rate, open sourced models are expected to surpass cloud models within a couple years based on a study I read a couple days ago. Looking back at chatGPT and claude a couple years ago, very small Qwen models are basically equal in coding to what those cloud based models could do then. Also factoring in scaling laws, a 9b going to 18b is roughly a 40% increase, whereas 18b to 35b is 20%, I expect there wi…

That makes no sense, though, and reeks of extrapolating a trend way beyond the conditions in which it is valid. The simple truth is, cloud models are always going to be strictly superior to open ones, simply because cloud model vendors can run those same open models too . And they still retain economies of scale and efficiency that operating large data centers full of specialized hardware, so at the very least they c…

I don't think the real-world evidence supports your argument... OpenAI and Anthropic have all of those advantages today, and Chinese models are reaching the same level. Clearly, the Chinese labs are doing something very right that is not directly related to infinite money.
Post reply on HN