If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…
Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
281–290 of 491 posts
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#282Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#283If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…
Fable works very well for me on a moderately large codebase. I have had to correct it a few times or point it on the right track, but given how much faster it is at programming than I am that's a very minor issue (and most of these errors are because I underspecified what I wanted in the prompt, I can only think of two cases where it was genuinely wrong... that's a lot better than me in my professional career). Code…
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#284Earlier quoted context omitted.
"Grammatically incomplete"?!
It writes in a bunch of annoying too-short sentence fragments, so, yes?
Even with my native language, which is definitely a lot less represented in the training data, the worst I encounter are phrasing mistakes, incorrect use of idioms, and invented words. You have to use some really badly tortured local model to get an LLM to produce incorrect grammar.
I do sometimes see Opus make typos, which is entertaining, but again, not a grammar issue.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#285If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#286If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…
Having tested K3, Qwen 3.8 max preview, Fable and Sol for the past few days, I partially agree. Don’t trust the benchmarks, and the Chinese models really are slow and token-inefficient. However they do seem very close to SOTA : I’d say roughly equal to previous gen (Opus 4.8, GPT 5.5). It’s yet another silly benchmark, but compare them here: https://senko.net/vibecode-bench/ I also had K3, Qwen3.8 and Fable (using Ki…
Its hard to compare that one to my regular sub of anthropic (107€), but everyone always says how much cheaper it is so i thought it was worth a try.
1.5 days later ive went through more then 50% of my weekly usage and decided to renew the regular anthropic sub too.
The API Pricing is definitely cheaper, but anthropics subscription budget seems to equalize that advantage right now. also kimi code feels like claude code from ... september 2025
also - considering how they announced they'd temporarily close subscriptions to make sure they can service customers... i was slightly surprised that there was no issue subscribing, less then 1 day after that announcement. Makes you think if that was just a marketing stunt
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#287Earlier quoted context omitted.
[flagged]
"there is literally not a single major industry where the US is not a leading player." Oof you were doing so strong until you threw this one out. Easy counterexample (and there are so many more than this). Clothing. Unfortunately, most people aren't buying MIUSA Selvedge Denim, PNW boots. I'm pretty sure that MIUSA clothing is like, 3% or less of all clothing sold in the USA.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#288Why SoTA (uppercase “T”) instead of SotA (lowercase “T”) ? “State of [T]he Art” versus “State of [t]he Art”. If not SotA then at least SOTA, which is more accurate.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#289If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…
You got baited by bad sampling settings. It's exactly the opposite. Go turn on min_p once it's available post July 27th and most of the problems you describe will go away.
This seems both arrogantly dismissive ("you are holding it wrong") and incorrect.
Either the OP is using Kimi K3 on Moonshot where is is presumable set correctly (K3 isn't available elsewhere yet), or they are using Kimi K2.x and there has been plenty of time to experiment with this.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#290If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…
We run an evaluation that only compares models in open-ended multi-agent environments where agents affect each other, primarily testing writing code. It's designed to be less vulnerable because there's no solution set, and it's been pretty effective and tends to rank Chinese models lower than their advertised model cards (relative to US models). Kimi K3 is a bit of an exception there -- it truly is a near-frontier model. But it's so slow. Ironically, Muse Spark 1.1 is one of the strongest models we've tested after Fable and Sol while also leading the cost efficiency curve. Big turnaround from Llama 4.
Data at https://gertlabs.com/rankings