Wow, if the benchmarks checkout with the vibes, this could almost be like a Deepseek moment with Chinese AI now being neck and neck with SOTA US lab made models
With the previous generation? Yes. With 10T mythos-level models? Not even close.
Kimi K2.6: Advancing open-source coding
21–30 of 394 posts
According to the benchmarks, you are wrong. It is on track and slightly above some sota. Just the benchmarks speaking there, they can be/are gamed by all big model labs including domestic.
Re: Kimi K2.6: Advancing open-source coding
#22I pray the benchmark figures are true so I can stop paying Anthropic after screwing me over this quarter by dumbing down their models, making usage quotas ridiculously small, and demanding KYC paperwork.
Anthropic has done horrible PR and investors should be livid.
Re: Kimi K2.6: Advancing open-source coding
#23I pray the benchmark figures are true so I can stop paying Anthropic after screwing me over this quarter by dumbing down their models, making usage quotas ridiculously small, and demanding KYC paperwork.
Anthropic has done horrible PR and investors should be livid.
My theory is they pushed retail off their systems to make room for their new corporate fat cat clients. In which case, they'll do just fine.
Re: Kimi K2.6: Advancing open-source coding
#24Wow, if the benchmarks checkout with the vibes, this could almost be like a Deepseek moment with Chinese AI now being neck and neck with SOTA US lab made models
With the previous generation? Yes. With 10T mythos-level models? Not even close.
10T? Impossible! They told us the training run was under 10^26 flops.
Re: Kimi K2.6: Advancing open-source coding
#25Re: Kimi K2.6: Advancing open-source coding
#26I've always been surprised Kimi doesn't get more attention than it does. It's always stood out to me in terms of creativity, quality... has been my favorite model for awhile (but I'm far from an authority)
Re: Kimi K2.6: Advancing open-source coding
#27There is some humor in the fact that china (of all countries) is pioneering possibly the world's most important tech via open source, while we (US) are doing the exact opposite.
Re: Kimi K2.6: Advancing open-source coding
#28Re: Kimi K2.6: Advancing open-source coding
#29The choice of example task for Long-Horizon Coding is a bit spooky if you squint, since it's nearing the territory of LLMs improving themselves.
Re: Kimi K2.6: Advancing open-source coding
#30If the benchmarks are private, how do we reproduce the results? I looked up the Humanity's Last Exam (https://agi.safe.ai/) this model uses and I can't seem to access it.