As always, benchmarks rarely paint the whole picture. It also seems like this article is somewhat biased, eg when Fable and Kimi are close but Fable wins it’s “dead heat”, but when Kimi wins it’s “Kimi wins”. GPT 5.6 seems to be missing as well. I am really eager to give Kimi K3 a try, but I’ll reserve my judgement until I’ve worked with it for at least a few days.
The apparent bias may be explainable as it’s not remarkable for OpenAI or Anthropic to be slightly ahead. It _is_ remarkable for an open weights model to be better than the closed models from the trillion dollar (allegedly) companies.
Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
211–220 of 491 posts
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#212Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#213I enjoy all the fun of these new models as much as the next guy but I truly don’t see a circumstance in the near future where my $200 a month with the frontier labs doesn’t get me more than enough consumption of what I need. Local models, chinese models, etc are all very fun weekend projects to tinker with but until something changes (entirely possible!) with how much you get with one of the subscriptions I just don’…
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#214Earlier quoted context omitted.
The best part is using harnesses like reasonix or whale make cache hit at a rate close to 98%, making requests converge to practically free. And that's with unsubsidized American providers like cloudflare or Digital Ocean.
How can you be hitting cache on what I think are novel LLM prompts …
Deepseek afaik has a novel architecture that is somewhat forgiving of cache shifts. I’ve been getting +90% cache hits with Zed’s agent and it’s not doing anything special regarding caching afaik.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#215I will accept a 5% drop in benchmarks for a model that talks to me like a human.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#216Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#217I'm skeptical. According to arena.ai, Fable 5 dominates almost every category : https://arena.ai/leaderboard Kimi K3 has an edge in WebDev but struggles to reach top 10 in many other categories.
In my experience, Fable is not even close to Sol 5.6 High (not even the max tier) for coding. 1) it's substantially slower. 2) it's substantially more expensive. 3) it's code is considerably worse. It's a joke when you consider what you get for what you pay for.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#218I will accept a 5% drop in benchmarks for a model that talks to me like a human.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#219I will accept a 5% drop in benchmarks for a model that talks to me like a human.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#220Earlier quoted context omitted.
To the public. Mythos has been in active use for quite a while.
It makes no good sense to evaluate models we don't have access to. For all we know K3 was competitive back then too. Or maybe there's a K4 in the works that blows everything out of the water. Who knows and who cares. There's no way for us to compare
Mythos was withheld because of the threat to security and/or marketing stunt (depending on your leaning), I don't see what benefit there could be for not releasing K3 immediately.