Live data from Hacker News

Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

fireworks.ai

281–290 of 491 posts

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#281

If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…

[deleted]

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#283

If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…

Fable works very well for me on a moderately large codebase. I have had to correct it a few times or point it on the right track, but given how much faster it is at programming than I am that's a very minor issue (and most of these errors are because I underspecified what I wanted in the prompt, I can only think of two cases where it was genuinely wrong... that's a lot better than me in my professional career). Code…

Fable is the clearly best when you have to do real coding.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#284
post #279

Earlier quoted context omitted.

"Grammatically incomplete"?!

It writes in a bunch of annoying too-short sentence fragments, so, yes?

Sentences being too short and annoying to read is a stylistic grievance, not a grammar one. The examples were edited in after my original reply, but even then, I see no actual grammatical mistakes there.

Even with my native language, which is definitely a lot less represented in the training data, the worst I encounter are phrasing mistakes, incorrect use of idioms, and invented words. You have to use some really badly tortured local model to get an LLM to produce incorrect grammar.

I do sometimes see Opus make typos, which is entertaining, but again, not a grammar issue.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#285

If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…

All you've said needs the qualifier - "for now!" Look at the trend line. It's clear that if they're not yet at the level of being "good enough" for coding, they will be soon. Sensationalist headlines aside, we all need to be preparing for a world where open models can do pretty much any software tasks you need them to.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#286
post #276

If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…

Having tested K3, Qwen 3.8 max preview, Fable and Sol for the past few days, I partially agree. Don’t trust the benchmarks, and the Chinese models really are slow and token-inefficient. However they do seem very close to SOTA : I’d say roughly equal to previous gen (Opus 4.8, GPT 5.5). It’s yet another silly benchmark, but compare them here: https://senko.net/vibecode-bench/ I also had K3, Qwen3.8 and Fable (using Ki…

with all the hype around kimi i actually renewed my subscription (last time i tried it out was on 2.5 release) - just the 40€ one.

Its hard to compare that one to my regular sub of anthropic (107€), but everyone always says how much cheaper it is so i thought it was worth a try.

1.5 days later ive went through more then 50% of my weekly usage and decided to renew the regular anthropic sub too.

The API Pricing is definitely cheaper, but anthropics subscription budget seems to equalize that advantage right now. also kimi code feels like claude code from ... september 2025

also - considering how they announced they'd temporarily close subscriptions to make sure they can service customers... i was slightly surprised that there was no issue subscribing, less then 1 day after that announcement. Makes you think if that was just a marketing stunt

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#287

Earlier quoted context omitted.

[flagged]

"there is literally not a single major industry where the US is not a leading player." Oof you were doing so strong until you threw this one out. Easy counterexample (and there are so many more than this). Clothing. Unfortunately, most people aren't buying MIUSA Selvedge Denim, PNW boots. I'm pretty sure that MIUSA clothing is like, 3% or less of all clothing sold in the USA.

The US is the second largest textile exporter in the world. Sources: https://ncto.org/facts-figures/us-textile-industry/ and https://www.trade.gov/selectusa-textiles-industry It's true that a significant amount of that is fabric and fibers, and most of the clothes sold in the US are not actually sewed here. But on the other hand, US companies also show up in other places in the clothing supply chain as well - American fashion design and marketing contribute a lot of the final value of many clothing items sold here and overseas.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#289

If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…

You got baited by bad sampling settings. It's exactly the opposite. Go turn on min_p once it's available post July 27th and most of the problems you describe will go away.

> Go turn on min_p once it's available post July 27th and most of the problems you describe will go away.

This seems both arrogantly dismissive ("you are holding it wrong") and incorrect.

Either the OP is using Kimi K3 on Moonshot where is is presumable set correctly (K3 isn't available elsewhere yet), or they are using Kimi K2.x and there has been plenty of time to experiment with this.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#290

If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…

Yes, they are all benchmaxxed, but the question is how benchmaxxed they are relative to each other.

We run an evaluation that only compares models in open-ended multi-agent environments where agents affect each other, primarily testing writing code. It's designed to be less vulnerable because there's no solution set, and it's been pretty effective and tends to rank Chinese models lower than their advertised model cards (relative to US models). Kimi K3 is a bit of an exception there -- it truly is a near-frontier model. But it's so slow. Ironically, Muse Spark 1.1 is one of the strongest models we've tested after Fable and Sol while also leading the cost efficiency curve. Big turnaround from Llama 4.

Data at https://gertlabs.com/rankings

Post reply on HN