Live data from Hacker News

Kimi K2.6: Advancing open-source coding

kimi.com

291–300 of 394 posts

Re: Kimi K2.6: Advancing open-source coding

#291
post #39

Earlier quoted context omitted.

While I'm skeptical of any "beats opus" claims (many were said, none turned out to be true), I still think it's insane that we can now run close-to-SotA models locally on ~100k worth of hardware, for a small team, and be 100% sure that the data stays local. Should be a no-brainer for teams that work in areas where privacy matters.

I think this one is only about 600GB VRAM usage, so it could fit on two mac studios with 512GB vram each. That would have costed (albeit no longer available) something like less than 20k.

That's just throwing money away. The performance with large context would have been unusable especially if you need to serve more then a single person.

Re: Kimi K2.6: Advancing open-source coding

#292
post #156

Earlier quoted context omitted.

The reason regular-capitalism worked is that all production used to depend on workers bottlenecking the free flow of capital by demanding salaries in exchange for their labor. Now that we've removed that obstacle, capitalism demands workers seize the means of production in order to maintain the status quo. Hence, supercapitalism.

workers seizing the means of production is by definition socialism and not capitalism though, that's the whole idea behind socialism

You miss the point: we advertise the change as workers becoming part of the owner class and realizing all of the economic gains of their work, thus supercapitalism. Don't use the "s" or "c" words.

Re: Kimi K2.6: Advancing open-source coding

#293
post #45

Earlier quoted context omitted.

What's the privacy/data security like? I can't find that on that page. Edit: found it. > We may use your Content to operate, maintain, improve, and develop the Services, to comply with legal obligations, to enforce our policies, and to ensure security. You may opt out of allowing your Content to be used for model improvement and research purposes by contacting us at membership@moonshot.ai. We will honor your choice i…

You really rely on ToS from Anthropic/OpenAI to know if they use your prompts or not? It's on their servers, why wouldn't they use our data?

Antropic and OpenAI are used by US businesses and government and they are audited and under contracts.

If it's discovered they trained on data they shouldn't have had it will be the end of their business.

On the other hand, good luck suing a Chinese company.

Re: Kimi K2.6: Advancing open-source coding

#294
post #231

Earlier quoted context omitted.

I think you should worry more about NSA, FBI, ICE and other 3 letter US agencies monitoring your sessions

There's nothing anyone can do about state-level espionage anywhere, using any cloud-hosted service. That being said, there is a very big difference between the legal situation in the United States vs. China. Chinese internet companies are required to have CPC interaction and since the rule of law does not strictly exist in China, the state can compel surveillance cooperation regardless of what might be written down.…

> the rule of law does not strictly exist in China, the state can compel surveillance cooperation regardless of what might be written down

While I agree that China is obviously worse in this regard, it's naive to claim this is unique to China, when literally a couple of months ago the US got into a fight with Anthropic about them not removing safeguards which were already just enforcing the letter of the law.

Re: Kimi K2.6: Advancing open-source coding

#295

Early benchmarks show tremendous improvement over Kimi K2 Thinking, which didn't perform well on our benchmarks (and we do use best available quantization). Kimi K2.6 is currently the top open weights model in one-shot coding reasoning, a little better than GLM 5.1, and still a strong contender against SOTA models from ~3 months ago (comparable to Gemini 3.1 Pro Preview). Agentic tests are still running, check back t…

How would K2.6 compare to Sonnet 4.6 both price and performance wise?

Re: Kimi K2.6: Advancing open-source coding

#296

Early benchmarks show tremendous improvement over Kimi K2 Thinking, which didn't perform well on our benchmarks (and we do use best available quantization). Kimi K2.6 is currently the top open weights model in one-shot coding reasoning, a little better than GLM 5.1, and still a strong contender against SOTA models from ~3 months ago (comparable to Gemini 3.1 Pro Preview). Agentic tests are still running, check back t…

Surprised to see such variance per language

It's interesting; I can only speculate as to the underlying reason. When given enough time, models outperform in Rust/C++ in longer agentic tasks, and actually perform worst in Python. For tasks that aren't judged on code speed. https://gertlabs.com/?mode=agentic_coding

Re: Kimi K2.6: Advancing open-source coding

#297
post #295

Early benchmarks show tremendous improvement over Kimi K2 Thinking, which didn't perform well on our benchmarks (and we do use best available quantization). Kimi K2.6 is currently the top open weights model in one-shot coding reasoning, a little better than GLM 5.1, and still a strong contender against SOTA models from ~3 months ago (comparable to Gemini 3.1 Pro Preview). Agentic tests are still running, check back t…

How would K2.6 compare to Sonnet 4.6 both price and performance wise?

In terms of raw token cost, I've seen a couple providers at (all prices in terms of Mtok) $0.95 input/$0.15 cache input/$5 output vs $3 input/$15 output for sonnet.

Task prices of courses will be more interesting - a dumber model may use more tokens to get to the same goal.

Re: Kimi K2.6: Advancing open-source coding

#298
post #228

Earlier quoted context omitted.

Maybe this can help https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle/

It doesn't, I get that it's _a_ benchmark. It's just not a good or insightful one, and having it posted so often on HN feels like low quality spam at this point

The issue is that benchmarks that look insightful will end up being gamed by labs quickly (Goodharts law)

The best LLM benchmarks test around the margins of those behaviors, tasks that are difficult and correlate with usefulness while being removed enough to stay unpolluted

Re: Kimi K2.6: Advancing open-source coding

#299

Earlier quoted context omitted.

At this point drawing these Pelicans must be in the training data sets.

not if I can help it! https://github.com/scosman/pelicans_riding_bicycles

These pelicans are clearly indicative of good RL training algorithms.

Re: Kimi K2.6: Advancing open-source coding

#300

There is some humor in the fact that china (of all countries) is pioneering possibly the world's most important tech via open source, while we (US) are doing the exact opposite.

I'm genuinely so grateful for them $200/m minimum to use Claude would bankrupt my country's white collar labor market

I would really appreciate a response because I'm sure you know that Anthropic has at least two lower priced tiers before the $200/m one, so I assume the $200/m tier is necessary because you use it heavily?

Now given that the $200/m Tier is the most heavily (I believe at 20x?) subsidized tier, How or what are you using instead that achieves comparable good enough performance for a fraction of the price? I've heard GLM 5.1 from z.ai but it's not comparable to Opus, not even close - really interested!

Post reply on HN