Anti-China: K3 is propaganda and benchmaxxed, no matter how anthropic and openai reactor for these, it just a smoke signal. Pro-China: K3 is good choice for better and affordable choice to smash down the Big three ruling.
China-ambivalent: Open Models are good, Closed Models are bad. Not centralizing power in a few big American companies is good. China is who's building this right now, so we're aligned for now.
Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
371–380 of 491 posts
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#372If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…
Anthropic/OpenAI were touting PhD-level intelligence three years ago. And they’re still shipping models that aren’t smart enough to realize things such as the need to drive the car to the car wash (because they hadn’t yet hill-climbed that particular brain-teaser).
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#373Earlier quoted context omitted.
One benefit of an open source one is that you can, as a large corporation, run it "locally" within your own data center. Even fine tune it.
How big is this market, self-hosting a model that requires 64 GPUs, H100 or better, with good interconnects between nodes? I suspect the overlap of those that can afford it, and those that have the talent to manage it, is a fairly thin slice of the Venn diagram. Even the large corps are gonna be getting it from the inference vendors, or more likely Bedrock and friends.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#374Very interesting. They test Kimi K3 and Fable on a set of approx 1000 tasks grouped into 5 areas (SWE, Legal, etc). They put a router model in front that predicts whether Kimi or Fable is going to give a better cost for a correct result. (They believe that ultimately such a router model should be continuously trained on your own workloads so it makes the best decisions for you). Their router chose Kimi the majority o…
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#375Earlier quoted context omitted.
I strictly prefer when models ignore any human quirks in my responses. Claude trying to be your friend, saying LOL to your jokes is ridiculous and frankly, harmful
Claude (Opus 4.8) recently told me: > I’d ask you to drop the abuse; (...) if it continues I’ll end the conversation. After I'd used a couple of expletives. And yes it will emit a token. This is truly dystopian. It is NOT a person. What a response. I still cant believe it.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#376Earlier quoted context omitted.
> Oracle routing is a method for measuring the best theoretical performance by running the task through each model and then picking the cheapest correct option (the cost/performance ceiling). Their "router" is an oracle reference point where they choose the lower cost model after running both and therefore knowing who passed the test. The cost savings part is only Fireworks theorizing what would happen if an equivale…
> running the task through each model and then picking the cheapest correct option What system knows what the correct option is and how does it know it?
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#377Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#378Anti-China: K3 is propaganda and benchmaxxed, no matter how anthropic and openai reactor for these, it just a smoke signal. Pro-China: K3 is good choice for better and affordable choice to smash down the Big three ruling.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#379I'm skeptical. According to arena.ai, Fable 5 dominates almost every category : https://arena.ai/leaderboard Kimi K3 has an edge in WebDev but struggles to reach top 10 in many other categories.
In my experience, Fable is not even close to Sol 5.6 High (not even the max tier) for coding. 1) it's substantially slower. 2) it's substantially more expensive. 3) it's code is considerably worse. It's a joke when you consider what you get for what you pay for.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#380Earlier quoted context omitted.
Having tested K3, Qwen 3.8 max preview, Fable and Sol for the past few days, I partially agree. Don’t trust the benchmarks, and the Chinese models really are slow and token-inefficient. However they do seem very close to SOTA : I’d say roughly equal to previous gen (Opus 4.8, GPT 5.5). It’s yet another silly benchmark, but compare them here: https://senko.net/vibecode-bench/ I also had K3, Qwen3.8 and Fable (using Ki…
Could you check how many tokens the different models spent to do the task? With all the comments about Kimi K3 Not being token efficient I'm curious if your test confirms that
* Kimi K3: 9532k input (9172k cached), 114k output - cost $5.5
* Qwen 3.8 Max: 18020k input (17836k cached), 114k output - cost $6.3
* Fable: ~14m input (all cached??), 196k output - cost $30
Correction on my earlier post, Kimi was through Pi, not Kimi Code. For Qwen I used Qwen Code and for Fable I used Claude Code.
Not sure wtf is going on with the Fable stats (a lot tokens, virtuall all of them were cached - I guess heavy system prompt?) but both claude code stats and ccusage tool output match.