Live data from Hacker News

Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

thinkpol.ca

201–210 of 235 posts

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#201

Earlier quoted context omitted.

DeepSeek and other Chinese model makers are massively accelerating progress in AI not slowing it down. They're the only ones who still come up with real technical innovations while the proprietary model makers are stagnating.

Can you name some tangible AI idea that came out of Chinese labs? I can name thousands that came out western universities. I see a lot of rhetoric that only the Chinese labs are contributing to AI while companies like Google and Microsoft are still pulishing their research. Unfortunately the domain of scientific papers is cluttered with AI slop but still occasional serious paper that i find are from western labs part…

Oh please https://github.com/deepseek-ai

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#202

Earlier quoted context omitted.

You never get "the same" Steph Curry, he might be tired, annoyed by a fan, getting older... but if he and I were to throw 100 3-pointers, we could all correctly guess who will perform better.

Good point. But I use Codex and Claude daily (work and hobby respectively). And there are days where one or the other just seems to have gotten up on the wrong side of the bed. Or is just being lazy. Or is suddenly super-powered do everything including what i asked it not to. (To be fair, the same thing happens with myself. :/) I am convinced that if I was bench-marking, I would be convinced these are different model…

That's also fair, Anthropic lobotomized their services a couple of times already. One week, you are in awe that the tools figure out everything, explain everything, consider everything, produce a clean fix... next week, they are completely useless.

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#203

Earlier quoted context omitted.

That makes no sense, though, and reeks of extrapolating a trend way beyond the conditions in which it is valid. The simple truth is, cloud models are always going to be strictly superior to open ones, simply because cloud model vendors can run those same open models too . And they still retain economies of scale and efficiency that operating large data centers full of specialized hardware, so at the very least they c…

I don't think the real-world evidence supports your argument... OpenAI and Anthropic have all of those advantages today, and Chinese models are reaching the same level. Clearly, the Chinese labs are doing something very right that is not directly related to infinite money.

Doesn't change the argument. As long as the models are open, the big cloud providers have strict advantage, because even if some open model gets ahead, they can just serve it from their infra, and do it better than everyone else.

This proves the strict inequality in my claim is preserved, everything beyond that is just debating the size of their advantage.

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#204
post #123

I'm glad we're seeing a shift towards objectively scored tests. We've been doing this at scale at https://gertlabs.com/rankings , and although the author looks to be running unique one-off samples, it's not surprising to see how well Kimi K2.6 performed. Based on our testing, for coding especially, Kimi is within statistical uncertainty of MiMo V2.5 Pro for top open weights model, and performs much better with tools…

This may be objectively scored, but it is not an indication of anyone's coding capabilities. This test measures which model almost accidentally came up with the best strategy (against other bots). This is not representative of coding. You would need to test 100 or more of such puzzles, widely spread across the puzzle spectrum, to get an idea which model is best at finding strategies involving an English dictionary.

If you are referring to the parent post, yes, hard to draw conclusions from such a small sample size.

For our testing, we use hundreds of different environments across disciplines, and it seems to line up with subjective experience better than other benchmarks. We test coding, agentic coding, and non-coding reasoning in the environments.

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#206

At the current rate, open sourced models are expected to surpass cloud models within a couple years based on a study I read a couple days ago. Looking back at chatGPT and claude a couple years ago, very small Qwen models are basically equal in coding to what those cloud based models could do then. Also factoring in scaling laws, a 9b going to 18b is roughly a 40% increase, whereas 18b to 35b is 20%, I expect there wi…

$600/mo? Do you mean $600 as a one time purchase for life? I've never heard of any Adobe plan that expensive.

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#207
post #185

Earlier quoted context omitted.

DeepSeek and other Chinese model makers are massively accelerating progress in AI not slowing it down. They're the only ones who still come up with real technical innovations while the proprietary model makers are stagnating.

I'm as happy to see cheap open weight models any anyone is, and I'm in Europe and certainly not cheering the US on, but that's a bunch of unfounded hyperbole you just said.

You should read the research papers that come out with Deepseek releases. There is a reason why the first Deepseek release briefly caused existential panic.

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#208
post #185

Earlier quoted context omitted.

I'm as happy to see cheap open weight models any anyone is, and I'm in Europe and certainly not cheering the US on, but that's a bunch of unfounded hyperbole you just said.

You should read the research papers that come out with Deepseek releases. There is a reason why the first Deepseek release briefly caused existential panic.

I did not and am not inclined to invest the time to do so.

But I did read some second hand reports that what was new and exciting was that they found some really good performance optimizations. The thing about deekseek publishing this is that now everyone has this.

Or did I miss something?

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#209
I don't know much about the AI field, but it seems to me that trying to train any model to be all things to all people is a really dumb idea. It requires huge financial resources and is causing extreme shortages/market distortions in any resource used by an AI company - RAM, SSDs, data centers, etc.

In the real world, you don't hire a plumber and expect him to also do your landscaping, fix your car, and tailor your clothes. It would seem like a much better use of resources if I could download an app that specialized in shell, Python, and C coding for example, or maybe even that would be 3 apps that communicated. Maybe I could even run them on a regular machine with 16GB of RAM. I don't need one huge model that can do that and code in Fortran, COBOL, and Lisp.

As humans, we've done pretty well by specializing. I hope this gets explored more with smaller, focused AI models vs the current path of one model to rule them all that can only be run in a data center the size of a country.

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#210
post #209

I don't know much about the AI field, but it seems to me that trying to train any model to be all things to all people is a really dumb idea. It requires huge financial resources and is causing extreme shortages/market distortions in any resource used by an AI company - RAM, SSDs, data centers, etc. In the real world, you don't hire a plumber and expect him to also do your landscaping, fix your car, and tailor your c…

This is true unless it isn't basically.

People claim that finetuning is good because no model can be 'that' general since gpt3 and at every generation it's becoming less true

Post reply on HN