Live data from Hacker News

Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

thinkpol.ca

221–230 of 235 posts

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#221
post #51

Kimi is really good. I have been using Sonnet and others (DeepSeek, ChatGPT, MiniMax, Qwen) for my compiler/vm project and the Claude Pro plan is mostly unusable for any serious coding effort. So I use it in chat mode in the browser where it cannot needlessly read your entire project, and use Kimi on the OpenCode Go plan with pi. Kimi consistently exceeded Sonnet on the C+Python project. Never had to worry about it d…

>the Claude Pro plan is mostly unusable for any serious coding effort Why? Seems to go a giant the opinion of the masses who mostly use Claude Pro for serious coding.

Claude is opaque as regards token usage. So you might end up using your 5hr limit in 7-10 minutes using regular Sonnet. Meanwhile, OpenCode etc give you exact breakdown in terms of how many cached tokens used per session etc which you can use to estimate burn rate.

All these coding tools are extremely wasteful as far as resources are concerned. Almost designed to make you move to the next tier. You have to consciously restrict their scope all the time to make your plans last. Even with Kimi/MiniMax a 3-4 hour session often ends up with 50-70M cached reads. Not a small amount at all.

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#222
post #208

Earlier quoted context omitted.

You should read the research papers that come out with Deepseek releases. There is a reason why the first Deepseek release briefly caused existential panic.

I did not and am not inclined to invest the time to do so. But I did read some second hand reports that what was new and exciting was that they found some really good performance optimizations. The thing about deekseek publishing this is that now everyone has this. Or did I miss something?

> The thing about deekseek publishing this is that now everyone has this.

It sounds like you're agreeing with upstream comment then!

>> DeepSeek and other Chinese model makers are massively accelerating progress in AI not slowing it down

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#223

Earlier quoted context omitted.

I don't think the real-world evidence supports your argument... OpenAI and Anthropic have all of those advantages today, and Chinese models are reaching the same level. Clearly, the Chinese labs are doing something very right that is not directly related to infinite money.

Doesn't change the argument. As long as the models are open, the big cloud providers have strict advantage, because even if some open model gets ahead, they can just serve it from their infra, and do it better than everyone else. This proves the strict inequality in my claim is preserved, everything beyond that is just debating the size of their advantage.

> As long as the models are open, the big cloud providers have strict advantage, because even if some open model gets ahead, they can just serve it from their infra

Why would I want to use it, though? If, say, Anthropic were to serve a hypothetical Kimi K5.0 from their infra, seems like they'd keep their pricing where it is. If I can use that same model from kimi.com/Kimi Code, for less money (which seems like a safe bet in this scenario), then I wouldn't use Anthropic's offering. Even if Anthropic did lower prices, I doubt they'd be able to match kimi.com/Kimi Code.

> ... and do it better than everyone else.

Why would you assume this? That doesn't follow. "Better" has diminishing returns, and all of these companies have impressively scaled up already, and will continue to scale further in the coming years. And, regardless, I would absolutely use someone else's infra if it cost, say, 20% less, even if inference was a bit slower, or I hit rate limits more often (not usage limits, rate limits).

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#225
post #164

Earlier quoted context omitted.

At least in my experience switching from Claude Pro ($20/month) to Kimi 2.6 through ollama (also $20/month), I was almost always hitting my usage limit with Sonnet 4.6, but with ollama I haven't hit my usage a single time.

How many t/ps do you get with Kimi on Ollama?

I actually have no clue how I'd check, do you know how?

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#226
post #223

Earlier quoted context omitted.

Doesn't change the argument. As long as the models are open, the big cloud providers have strict advantage, because even if some open model gets ahead, they can just serve it from their infra, and do it better than everyone else. This proves the strict inequality in my claim is preserved, everything beyond that is just debating the size of their advantage.

> As long as the models are open, the big cloud providers have strict advantage, because even if some open model gets ahead, they can just serve it from their infra Why would I want to use it, though? If, say, Anthropic were to serve a hypothetical Kimi K5.0 from their infra, seems like they'd keep their pricing where it is. If I can use that same model from kimi.com/Kimi Code, for less money (which seems like a safe…

Isn't Amazon Bedrock doing something quite similar already? The obvious argument is "We have Kimi at home" i.e. no need to pay for Chinese-supplied APIs that might misuse your submitted data.

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#227
post #185

Earlier quoted context omitted.

DeepSeek and other Chinese model makers are massively accelerating progress in AI not slowing it down. They're the only ones who still come up with real technical innovations while the proprietary model makers are stagnating.

I'm as happy to see cheap open weight models any anyone is, and I'm in Europe and certainly not cheering the US on, but that's a bunch of unfounded hyperbole you just said.

>> DeepSeek and other Chinese model makers are massively accelerating progress in AI... They're the only ones who still come up with real technical innovations.

> that's a bunch of unfounded hyperbole you just said.

Calling the quote on top "unfounded hyperbole" betrays lack of knowledge and awareness about the subject. Keep in mind that when we talk about real technical innovations, we have in mind published research - not closed or hidden models, some of which we know only from hype but cannot even test. A cursory look at said research reveals more Chinese names than I can count.

Deepseek did introduce real technical innovations, they're in their papers, and there was plenty of talk about another "Sputnik moment" when their first model appeared. If you don't know what that means - it's the moment when the industry mobilizes to "accelerate progress" due to the unexpected appearance of strong competition.

There's a lot more to be said, but it wouldn't do much good to a person who's not following the trends.

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#228
post #101
post #63

Earlier quoted context omitted.

While this might be true I’m worried about the hardware side of things. What if you have a good enough model but the cloud model providers are better in procuring hardware for interference?

I personally believe that eventually manufacturers will want to sell more of their hardware and look for ways to sell hardware to consumers. isnt that situation quite similar to the days of early computers? I am for sure biased in hoping that will be the case

Perhaps for some very specific capabilities such as TTS, translation, voice recognition and so on. But for general intelligence models, better hardware just directly allows better models and that doesn't seem to be changing any time soon.

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#229
post #209

I don't know much about the AI field, but it seems to me that trying to train any model to be all things to all people is a really dumb idea. It requires huge financial resources and is causing extreme shortages/market distortions in any resource used by an AI company - RAM, SSDs, data centers, etc. In the real world, you don't hire a plumber and expect him to also do your landscaping, fix your car, and tailor your c…

> I could download an app that specialized in shell, Python, and C coding for example, or maybe even that would be 3 apps that communicated. Maybe I could even run them on a regular machine with 16GB of RAM. I don't need one huge model that can do that and code in Fortran, COBOL, and Lisp.

I would daresay for "coding tasks", you actually _want_ a model that can code "in all languages".

Sure, it might be that outdated language XYZ is really useless to you or the task you want, but being exposed to their limits, philosophy and concerns across environment, framework and organization, among other things, means for example you get insights of your problems from other areas and points of view.

That's afterall how we got Newtonian physics and calculus, right? A person studying physics someday noticed how the "math of the day" wasn't able to calculate some results without a lot of elbow grease. He then "found" the "missing math" and with it was able to generalize what at the time was considered a bunch of isolated phenomena into a cohesive corpus of knowledge.

So for example, I want my code to have mechanical sympathy like Fortran; well defined input/output interfaces, and not-interweaved control structures, like COBOL; stateless, side-effects-free business logic like Lisp.

Re: Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

#230
post #223

Earlier quoted context omitted.

Doesn't change the argument. As long as the models are open, the big cloud providers have strict advantage, because even if some open model gets ahead, they can just serve it from their infra, and do it better than everyone else. This proves the strict inequality in my claim is preserved, everything beyond that is just debating the size of their advantage.

> As long as the models are open, the big cloud providers have strict advantage, because even if some open model gets ahead, they can just serve it from their infra Why would I want to use it, though? If, say, Anthropic were to serve a hypothetical Kimi K5.0 from their infra, seems like they'd keep their pricing where it is. If I can use that same model from kimi.com/Kimi Code, for less money (which seems like a safe…

Cloud v. local is a different axis to secret v. open.

Claude is secret and cloud; Kimi on e.g. AWS is open and cloud; Kimi on your machine is open and local; If there are any closed and local models, I don't know what they are (Apple Intelligence, if I had to guess?)

I'd argue slightly differently from TeMPOraL: Cloud has advantages when the best models are the big ones. Right now this is so, but this may not always be the case. If we are in a world where the models stop improving at any point (for whatever reason) while hardware keeps getting better (it might or might not), then we may find the small cost benefit from operating at scale isn't worth the effort let alone the legal implications.

Unrelated, but for me the film called Kimi is higher in search rankings than the model is and oh wow we really do have a problem with the whole "finally the torment nexus" thing don't we.

Post reply on HN