Live data from Hacker News

Kimi K3: Open Frontier Intelligence

kimi.com

971–980 of 1001 posts

Re: Kimi K3: Open Frontier Intelligence

#971
post #321

Earlier quoted context omitted.

These benchmark numbers are insane. The days when China was 6 months behind are over? How are they doing this with so much less resources than the US??? I have so much respect for the researchers there

To summarise the full results table further down the page (which doesn't render on the page for me!): Kimi K3 beats each model (out of 35 benchmarks, excluding missing): vs Fable 5 : 12/35 (34%) (ties: 1) vs GPT 5.6 Sol : 19/34 (56%) (ties: 1) vs Opus 4.8 : 30/35 (86%) vs GPT 5.5 : 30/34 (88%) (ties: 2) vs GLM-5.2 : 19/19 (100%) Beats Opus 4.8 and GPT 5.5 on all programming and agentic programming benchmarks except T…

Astonishing. Considering none of the BigTech except Google (Microsoft, Apple, Meta, Amazon, Nvidia, SpaceX) have managed to challenge OpenAI & Anthropic frontier models, such achievements are scarcely believable.

Re: GLM-5.2: For a ~750b model, it holds up pretty good against models 3x its size (and ~10x the cost). Same goes for Tencent Hy3 and MiniMax M3, which almost match Opus 4.6 levels with ~295b params.

Re: Kimi K3: Open Frontier Intelligence

#972
post #429

Strictly dominates both Sonnet 5 and Opus 4.8 in both cost and performance: https://artificialanalysis.ai/models/comparisons/kimi-k3-vs-... https://artificialanalysis.ai/models/comparisons/kimi-k3-vs-...

Yup some here are in denial but what many said would happen did just happen. They're not "six months behind": the model is totally SOTA. Cheaper, faster and they don't just crush Sonnet 5 and Opus 4.8: on 6 of the 14 benchmarks they posted Kimi K3 is in front of Fable . Of course the shills are shifting their tone: this thread as devolved into "sure yup it's totally SOTA but it sucks because it'll use more tokens tha…

> I'll make a prediction

You just hope that BigLabs & BigTech doesn't gut out the talent from Chinese labs. They certainly have the money & impetus.

Re: Kimi K3: Open Frontier Intelligence

#973

Earlier quoted context omitted.

I'm pretty annoyed with how fast this feels. Wish MacOS was this fast launching things.

You're so right. When did Apple take the wrong turn? What caused to become this slow at launching simple apps? Gatekeeper? XProtect? Swift? SwiftUI?

Probably all of them? On modern versions, Gatekeeper still adds a bunch of milliseconds and possibly a RTT network request to Apple before the first launch of many apps; although I think Apple's binary checking is cached for 24 hrs or something.

Re: Kimi K3: Open Frontier Intelligence

#974
post #840

Earlier quoted context omitted.

Even the terminal works, this is madness. I was able to create a temp folder, echo hello > world, and then I could open the folder in finder and double-clicking the file opens it in a GUI text editor.

This is wild. Do you think it's really a one prompter? K2.5 had a linux frontend one-shot for display that was very good looking and smooth but very little of it had function. Should I just like, idk, stop using subscriptions and API this shiz?

It might one prompt, but modern LLMs only really shine in agentic loop harnesses (e.g. Kimi Work, Cowork, etc). The OS demo was produced with 1 prompt in an agentic loop in Kimi Work.

Re: Kimi K3: Open Frontier Intelligence

#975
post #605

Earlier quoted context omitted.

Mythos/Fable-class models have been around for at least 4 months internally in the US, and Kimi still isn't quite there, so I'd say the 6-months is still about right.

This is fair, with the caveat that we don't know for how long this model has been around internally in China either. So we can only go about appearance / releases.

At least ~2 weeks as the mystery model in TextArena turned out to be Kimi K3.

Re: Kimi K3: Open Frontier Intelligence

#976

Earlier quoted context omitted.

Have you used the same session for audit? so switched to K3? or used a new session for K3? K3 is sensitive to this, they wrote about it on their blog.

Do you have a link handy? I used same session, set it to k3 model. I’ll look at the blog but the result was so bad I am prepared to abandon. I should have saved the output. I think maybe it was a mistake to not use open router.

K3 is very different to K2, I wouldn't be surprised if there are different system prompts, parsing templates, etc; which confuse/poision the model's context.

Re: Kimi K3: Open Frontier Intelligence

#977

This might be the most impressive website generator demo I've seen: https://macos27.kimi.page Context from the person who prompted it: https://x.com/mweinbach/status/2077827886149439547

This is seriously impressive, it built a file system under the hood? I was able to create files, copy around and view/change it with terminal and UI. Absolutely insane

You can quit apps from the activity monitor. It really does appear to be simulating an OS underneath. I am blown away.

Re: Kimi K3: Open Frontier Intelligence

#978
post #705

Earlier quoted context omitted.

For me none of the models have managed to accurately replicate any given .png/.jpg image in SVG. I guess that requires both vision encoders and coding layers to work perfectly.

have you tried quiver.ai

Thanks, I will give it a shot.

Re: Kimi K3: Open Frontier Intelligence

#979
post #673

Earlier quoted context omitted.

It's a GTM strategy: https://try.works/why-chinese-ai-labs-went-open-and-will-rem...

>Open sourcing models is not a commercial risk because barely anyone can run them locally there are dozens of us!

Open weights are indeed a commercial risk because there's no shortage of companies that'll make a business out of running MaaS (model as a service), close & resell a fine-tuned or distilled variant (DeepSeeking of Llama / Cursorification of Kimi, for example).

Re: Kimi K3: Open Frontier Intelligence

#980

This might be the most impressive website generator demo I've seen: https://macos27.kimi.page Context from the person who prompted it: https://x.com/mweinbach/status/2077827886149439547

Crikey. It even bound Shift-Cmd-G to the filesystem picker thing.
Post reply on HN