Live data from Hacker News

Kimi K2.6: Advancing open-source coding

kimi.com

371–380 of 394 posts

Re: Kimi K2.6: Advancing open-source coding

#371

Early benchmarks show tremendous improvement over Kimi K2 Thinking, which didn't perform well on our benchmarks (and we do use best available quantization). Kimi K2.6 is currently the top open weights model in one-shot coding reasoning, a little better than GLM 5.1, and still a strong contender against SOTA models from ~3 months ago (comparable to Gemini 3.1 Pro Preview). Agentic tests are still running, check back t…

Surprised to see such variance per language

It makes sense when you consider LLMs don't generalize very well, so they're heavily dependent on how good (how varied as well as how high quality) the training data is

Re: Kimi K2.6: Advancing open-source coding

#372

Early benchmarks show tremendous improvement over Kimi K2 Thinking, which didn't perform well on our benchmarks (and we do use best available quantization). Kimi K2.6 is currently the top open weights model in one-shot coding reasoning, a little better than GLM 5.1, and still a strong contender against SOTA models from ~3 months ago (comparable to Gemini 3.1 Pro Preview). Agentic tests are still running, check back t…

Cool website. I don't understand enough about the various benchmarks or how they're done to judge whether or not anything is accurate, but I love the layout and features especially the spectator feature which is pretty cool. One thing, I saw the "Market simulator" spectator feature but didn't see a corresponding benchmark for that. Is it "Finance" or "Betting" or "Trading"?

Thanks -- that one is categorized under Trading/Financial, whereas betting is reserved for games like Pot Limit Omaha Hilo.

That's a good idea for a feature request, including the tags for the spectatable demo games.

Re: Kimi K2.6: Advancing open-source coding

#373

Earlier quoted context omitted.

wait why compare 2.6 to 2 instead of to 2.5?

Good question. We missed that release entirely. Our automated model checker only went live 2 months ago so they were manually curated prior to that. I'm adding it now. It'll be live in ~12 hours.

Update: Kimi K2.5 one-shot results are live. It wasn't a noteworthy release compared to K2.6: https://gertlabs.com/?mode=oneshot_coding

Re: Kimi K2.6: Advancing open-source coding

#374
post #329
post #255

Earlier quoted context omitted.

I’m not sure what A/B test you’re part of but on Claude Code Pro, I hit every single one of my quotas without exception. If you analyze/process images it’s even worse: I hit rate limits first and if I use separate sessions, I hit my quotas too. I use up so many tokens that Jensen should hire me.

I specifically stated "chat" and "not sure about coding usage" but you're saying "Claude Code Pro".

You're right, I missed that part.

Re: Kimi K2.6: Advancing open-source coding

#375

Earlier quoted context omitted.

Surprised to see such variance per language

It makes sense when you consider LLMs don't generalize very well, so they're heavily dependent on how good (how varied as well as how high quality) the training data is

Well it might explain why pro-Claude vs pro-Codex people keep talking past each other on this forum. I see people all the time assuming that anybody who likes Codex must be some sort of bot because of their own biases, but I work almost exclusively in Rust and find Codex extremely competent (and a much better overall engineer), don't trust Claude/Opus at all... but I see in this bench it scores lower on TypeScript etc. than Opus does.

Re: Kimi K2.6: Advancing open-source coding

#376

Earlier quoted context omitted.

A huge dual socket Epyc system used to be able to get to 1TB without difficulty. 16 dimms of 64gb each. Doable for ~$3000. With considerable memory bandwidth. Our hope these days seems to be that maybe perhaps possibly High Bandwidth Flash works out. Instead of 4, 8, or maybe more for some highest end drives, having many many many dozens of channels of flash. Ideally that can be very very near to the inference. PCIe…

You can't buy 16 64gb dimms for $3000. Go shop memory prices again. But yes an old epyc can run this with no GPU at reasonable speed and if you throw a few GPUs you can get very manageable speed. I run this at home on an old system PCIe4, slow 2400mhz ddr4 ram and still getting about 13tk/sec

The "used to be able" in the first sentence is what I thought made it clear that I was talking past tense. The cost is indeed 10x what it was.

Re: Kimi K2.6: Advancing open-source coding

#377

Earlier quoted context omitted.

This perspective is pretty interesting: https://federicocarrone.com/articles/china-commoditizing-the...

Summary: they want to commoditize the complement which means that Western "knowledge work" is the complement to Chinese manufacturing, and they want to turn the knowledge work into a low priced commodity via open llm models. I've heard this before, always accompanied by a several thousand word blog post. But frankly it sounds like it's overcomplicating the issue. Why would you try to turn something into a commodity w…

That's a fair point. That probably makes more sense, especially when viewed from a company-specific perspective. Each individual actor probably has much more to gain by trying to actually compete than by trying to commoditize the complement.

If viewed from a national perspective, then the decision calculus could get more confusing. I can imagine that commoditizing LLMs might cost substantially less than trying to be a leader in the space. Of course, there is also less to gain in commoditizing LLMs versus being a leader.

I'm not sure, though, and you bring up good points.

Re: Kimi K2.6: Advancing open-source coding

#378
post #216

Earlier quoted context omitted.

The American models also censor a lot of scientific and political views though.

Can you be more specific?

Trump issued an EO against "woke AI" that allows them to directly influence how models respond

https://www.lawfaremedia.org/article/evaluating-the--woke-ai...

Re: Kimi K2.6: Advancing open-source coding

#379

Earlier quoted context omitted.

I think one of the motivations is undermining US companies. OpenAI and Anthropic are the two biggest players, and are American. Open weights models reduce the power those two big players have over the industry. If the Chinese companies tried to play by US rules and close-source their products then people would mostly use ChatGPT and Claude. So the Chinese companies don't make a ton of profit either way, but by releas…

I don't think so, it's just how things played out. Thanks to Meta, after llama leak and meta followed up with llama2 and llama3 that caused everyone else to follow up with open models, Stablediffusion, Mistral, Cohere, Microsoft phi, IBM granites, Nvidia Nemotrons, so the Chinese labs joined the fun too.

Stable Diffusion predates LLaMA

Re: Kimi K2.6: Advancing open-source coding

#380

Earlier quoted context omitted.

Wonder if stuff like this would affect it? https://github.com/p-e-w/heretic Guessing it probably would?

Neat project! I would be interested in a paper about this. I think the tricky part with this type of technology is that, this works if the training data was not curated. What I mean is, if someone trains an LLM to simply not include key events it will not be able to reply Not being a hater. This is neato!

In that case you can use either rag or fine-tuning. The entire premise of the Tiananmen Square argument is just Americans feeling inferior. I use Chinese models every day for work and my personal life, the model not knowing about this one historical event has had zero impact on me.
Post reply on HN