Live data from Hacker News

Kimi K2.6: Advancing open-source coding

kimi.com

341–350 of 394 posts

Re: Kimi K2.6: Advancing open-source coding

#341

Earlier quoted context omitted.

Except that didn't happen, and it's not a milestone. First, you are confusing share of electricity generation with the share of all energy. Electricity is only 21% of all energy. Natgas, oil and coal are crushing it in that remaining 79%. Second, the article is wrong, even for electricity. To their credit, Canary Media showed in their graph that this data is for electricity only. The data for March is not out yet. He…

> First, you are confusing share of electricity generation with the share of all energy. Think it was pretty obvious what I meant to all but the most pedantic, bud. But just to be clear, your issue here is that a think tank cited the same (notoriously anti-renewable Trump admin) government agency that you've cited multiple times yourself? That's what set off your spidey senses? Have you considered that this respected…

The EIA is where Ember gets its data from.

It's where everybody gets their data from. Because they have thousands of employees collecting data. These are professionals, like the people at BEA, HUD, NIST, etc.

Ember, on the other hand, is a "decarbonization" think tank. They don't have their own data. They don't have the staff for it. What they do is analyze/spin, and in this case, augment, the raw data that is published by EIA. How do they augment the EIA data? All they do is round it to the nearest 2 decimals. It's exact copy and paste for every month except the last two, where the data is just made up.

And this entire article was written based on the augmentations by Ember, yet Ember cites it as EIA data. So let's check back in July, when EIA data will be out, and Ember will use that exact data, rounding it to the nearest 2 decimals. Save that blog page!

Something to think about.

Re: Kimi K2.6: Advancing open-source coding

#342

Earlier quoted context omitted.

Could be! Simon wrote about that here though https://simonwillison.net/2025/Nov/13/training-for-pelicans-...

> If a model finally comes out that produces an excellent SVG of a pelican riding a bicycle you can bet I’m going to test it on all manner of creatures riding all sorts of transportation devices. This relies on the false premise that, if they would include it in their training dataset, it would be perfect. All they need to do is be good enough and better than the other, not perfect.

I'm not sure if we can have a "perfect" Pelican riding a bicycle. Like, I could probably commission a highly experienced artist to draw one and I don't think it would be perfect. The legs would probably have to be too long, or pedals oddly placed, or handles strange, or wings with hands.

Based on the one Simon commented though, I'd say we're in decent territory to try the latter part of his hypothesis.

Re: Kimi K2.6: Advancing open-source coding

#343

Earlier quoted context omitted.

> First, you are confusing share of electricity generation with the share of all energy. Think it was pretty obvious what I meant to all but the most pedantic, bud. But just to be clear, your issue here is that a think tank cited the same (notoriously anti-renewable Trump admin) government agency that you've cited multiple times yourself? That's what set off your spidey senses? Have you considered that this respected…

The EIA is where Ember gets its data from. It's where everybody gets their data from. Because they have thousands of employees collecting data. These are professionals, like the people at BEA, HUD, NIST, etc. Ember, on the other hand, is a "decarbonization" think tank. They don't have their own data. They don't have the staff for it. What they do is analyze/spin, and in this case, augment, the raw data that is publis…

I feel like I shouldn't have to be finding this info for you since it was right there in the links you already sent, but:

> Annual electricity generation and net imports are taken from the EIA.

> Monthly generation and imports are taken from the EIA. The EIA reports monthly generation data in two separate datasets: Monthly data for all 50 states and monthly data for the lower 48 states (excludes Hawaii and Alaska). Data for all 50 states is reported on a 3 month lag whereas data for the lower 48 states is reported without lag. Missing months from the data for all 50 states is estimated using the recent changes observed in data from the lower 48 dataset.*

Page 89: https://ember-energy.org/app/uploads/2024/05/Ember-Electrici...

There are two different EIA datasets.

Re: Kimi K2.6: Advancing open-source coding

#344

There is some humor in the fact that china (of all countries) is pioneering possibly the world's most important tech via open source, while we (US) are doing the exact opposite.

This perspective is pretty interesting: https://federicocarrone.com/articles/china-commoditizing-the...

Summary: they want to commoditize the complement which means that Western "knowledge work" is the complement to Chinese manufacturing, and they want to turn the knowledge work into a low priced commodity via open llm models.

I've heard this before, always accompanied by a several thousand word blog post. But frankly it sounds like it's overcomplicating the issue. Why would you try to turn something into a commodity when instead you could turn it into a trillion dollar industry and win?

The goal has always been clear:

1. Release open models to get your name out

2. Then once you feel you have name recognition release even stronger models but keep them proprietary. Qwen is clearly at this phase.

3. Keep releasing open models because it's good publicity but never your SOTA models (e.g. Google's Gemma).

Re: Kimi K2.6: Advancing open-source coding

#345

Earlier quoted context omitted.

Aligning a model in a way that causes it to refuse requests to produce propaganda for one country, but not for another country is what? Is there some functionally equivalent word to censorship you'd like to use because of you're naive enough to think US corporations would not self-censor but Chinese corporations would? - Also, you are invested the goalpost of "no matter how hard you try", I don't find it interesting…

Bias and censorship are not identical. The subject of this thread is censorship, not bias. Besides, why do you want a model to produce propaganda? Surely you have better things to do.

"Surely you have better things to do."

I certainly gave the hypothetical reader too much credit.

Re: Kimi K2.6: Advancing open-source coding

#346
post #53

Accessed via OpenRouter, this one decided to wrap the SVG pelican in HTML with controls for the animation speed: https://gisthost.github.io/?ecaad98efe0f747e27bc0e0ebc669e94... Transcript and HTML here: https://gist.github.com/simonw/ecaad98efe0f747e27bc0e0ebc669...

It looks like a drunk pelican rolling downhill on its bicycle

Re: Kimi K2.6: Advancing open-source coding

#347

Early benchmarks show tremendous improvement over Kimi K2 Thinking, which didn't perform well on our benchmarks (and we do use best available quantization). Kimi K2.6 is currently the top open weights model in one-shot coding reasoning, a little better than GLM 5.1, and still a strong contender against SOTA models from ~3 months ago (comparable to Gemini 3.1 Pro Preview). Agentic tests are still running, check back t…

Can you add Qwen 3.6 max to the leaderboard?

We will as soon as API access is widely available. Once a model goes live, we typically have one-shot reasoning benchmarks up in ~8 hours and comprehensive agentic/combined benchmarks up after 24-48 hours. We're working on building relationships with each lab to have the results before launch.

Re: Kimi K2.6: Advancing open-source coding

#348

Early benchmarks show tremendous improvement over Kimi K2 Thinking, which didn't perform well on our benchmarks (and we do use best available quantization). Kimi K2.6 is currently the top open weights model in one-shot coding reasoning, a little better than GLM 5.1, and still a strong contender against SOTA models from ~3 months ago (comparable to Gemini 3.1 Pro Preview). Agentic tests are still running, check back t…

wait why compare 2.6 to 2 instead of to 2.5?

Good question. We missed that release entirely. Our automated model checker only went live 2 months ago so they were manually curated prior to that. I'm adding it now. It'll be live in ~12 hours.

Re: Kimi K2.6: Advancing open-source coding

#349

Earlier quoted context omitted.

What does good even mean… I have no idea what a good “pelican on a bike” should look like. It’s a fun prompt because there is no good answers… at least so I thought.

Yeah that was exactly Simon's intent. https://simonwillison.net/2025/Nov/13/training-for-pelicans-...

There are countless examples of animals riding bicycles etc from Comic books I grew up with

It would always look goofy - by design, but it usually looked good.

Re: Kimi K2.6: Advancing open-source coding

#350

There is some humor in the fact that china (of all countries) is pioneering possibly the world's most important tech via open source, while we (US) are doing the exact opposite.

I think one of the motivations is undermining US companies. OpenAI and Anthropic are the two biggest players, and are American. Open weights models reduce the power those two big players have over the industry. If the Chinese companies tried to play by US rules and close-source their products then people would mostly use ChatGPT and Claude. So the Chinese companies don't make a ton of profit either way, but by releas…

This makes sense, but either ways, its a Big win for the consumers as these Chinese companies will keep the frontier labs' quality and prices honest.
Post reply on HN