Live data from Hacker News

Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

fireworks.ai

311–320 of 491 posts

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#311
post #88

Earlier quoted context omitted.

They share their methodology and results. I learned things about the relative strengths and weaknesses of Kimi and Fable I hadn’t seen anywhere else. Should being in the model hosting business disqualify them from sharing?

Doesn't disqualify them, but it may call into question their results seeing as they have a potential conflict of interest.

> a potential conflict of interest

That seems to be actually a simple direct interest.

And, there is the "my product is good" confirmation to the outside incentive, there is the (only potentially) conflicting truth incentive, but there are also internal mission studies needs - so that you do not research into the benchmarks just to show people that "your product is good".

> call into question their results

That's always valid, so the question becomes "how much", and at that point the important side is difficult to evaluate - especially because in a frequentistic, Bayesian context the quantities were just potential anyway ("This will increase the chances - yes, of course just the chances - of E by some amount").

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#312

If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…

So they aren’t as good as 2024 models? 2025?

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#313
post #245

Earlier quoted context omitted.

I believe the Chinese government is angling to destroy the western economy and rise from the ashes. Instead of a billion a day to bomb some buildings and bridges they're intentionally hamstringing the biggest concentration of speculation in history

I agree. The commenter you replied to makes it sound like the Chinese models are coming out of tiny startups with meager resources. It's really not a David vs Goliath story.

Is it? From what i could find Anthropic and OpenAI are hovering around 5000 employees where as Deepseek and Moonshot are more like 300. The funding/investments are similarly many times less.

I agree that 300 employees is not a small company but the overall outlook for Anthropic/OpenAI is not geat.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#314

Earlier quoted context omitted.

In long form tasks, across multiple harnesses (Claude Code, OpenCode, Kimi Code, ZCode) my cache rates are typically 96-99%. I don’t think that’s particularly out of the ordinary. Do people have different experiences with other harnesses? Which ones?

How much time do you spend on your setup vs getting a lot of stuff shipped by paying for fabel 5? For me, not using the best model is a huge opportunity cost since my company can afford it.

> How much time do you spend on your setup vs getting a lot of stuff shipped by paying for Fable 5?

Relatively speaking, quite little (and it's more interesting than some of the stuff I'm otherwise shipping), I mostly just explored to see what's available.

I also agree that non-SOTA models are risky, that's why I was shopping around for them as well - DeepSeek V4 Pro is cheap but unreliable, GLM 5.2 is around and maybe slightly past Sonnet quality (though their quotas are a bit of a problem), whereas Kimi K3 is a proper contender.

I did write a tool to manage 3rd party providers for Claude Code: https://ccode.kronis.dev/

However, in the end I figured out that for the terminal use cases OpenCode is really comfy (provider TUI solutions are okay).

For web/desktop based stuff I like the UI of Claude Code Desktop, though ZCode comes close too (which is surprising, they sorta came out of nowhere and are patching the thing weekly), meanwhile Kimi Code and OpenCode desktop/web offerings still aren't great, but are functional.

I actually did write a bit more about my experiences on my blog.

GLM 5.2 with their harness and coding plan: https://blog.kronis.dev/blog/z-ai-s-glm-5-2-is-a-great-model...

Kimi K3 with their Allegro 100 USD plan and harness: https://blog.kronis.dev/blog/kimi-k3-is-out-is-anthropic-don...

That exploration let me switch over to Kimi fully because the tone of Anthropic's models is insufferable: https://blog.kronis.dev/blog/ai-slop-is-a-self-inflicted-tra...

Not to spam too much, but maybe those experiences are useful to someone. Long story short, most harnesses are okay but shopping around a little bit is definitely a good idea, the same way how spending some time choosing a font isn't a bad thing if you'll stare at it for 8 hours a day. I'm also happy that I managed to find software that's okay to run and also a model whose tone I actually enjoy, that is still near-SOTA in performance and that I can throw tasks at it without worrying about whether the model is or isn't good enough at those.

I will admit that I'm probably slightly overspending by moving over fully to Kimi, there's probably a plateau for each kind of task and not everything needs SOTA models, but at least this way I don't have to think much about it.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#315
post #50

Earlier quoted context omitted.

Strawberries are a well known weakness of LLMs, as they have a hard time to count the numbers of "r"s in them. Maybe that's why, because they fear that weakness could be exploited somehow.

Strawberries aren't the weakness, the weakness is the tokenization of a prompt. Any word with multiple duplicate characters is going to be troublesome for LLMs.

I think the parent is aware and was joking.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#316

Genuine question: can these posts be paid to hype the open source models? If yes, what would be the purpose? On my work tasks, FastAPI Python and Springboot Java on a modern SaaS product, the only open model that can do tasks well and efficiently is Qwen3.7-Max. In all my experiments, both GLM-5.2 and Kimi are busy grepping around the codebase for ALMOST 70-80K tokens before writing anything and when they do it typic…

Have you tested Qwen3.8-Max yet?

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#317

If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…

[flagged]

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#318
post #289

Earlier quoted context omitted.

You got baited by bad sampling settings. It's exactly the opposite. Go turn on min_p once it's available post July 27th and most of the problems you describe will go away.

> Go turn on min_p once it's available post July 27th and most of the problems you describe will go away. This seems both arrogantly dismissive ("you are holding it wrong") and incorrect. Either the OP is using Kimi K3 on Moonshot where is is presumable set correctly (K3 isn't available elsewhere yet), or they are using Kimi K2.x and there has been plenty of time to experiment with this.

I've earned my right to be arrogantly dismissive since almost the entire field (including the Kimi and qwen team) doesn't know good sampling settings. This is because if you are truly "bitter lesson pilled" you don't think sampling is needed at all.

You can either not believe me and be wrong, or you can (after July 27th) turn on min_p or a better sampler (i.e. top-n-sigma if you got it running via llamacpp) and have it work even better. Up to you.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#319

Very interesting. They test Kimi K3 and Fable on a set of approx 1000 tasks grouped into 5 areas (SWE, Legal, etc). They put a router model in front that predicts whether Kimi or Fable is going to give a better cost for a correct result. (They believe that ultimately such a router model should be continuously trained on your own workloads so it makes the best decisions for you). Their router chose Kimi the majority o…

> Oracle routing is a method for measuring the best theoretical performance by running the task through each model and then picking the cheapest correct option (the cost/performance ceiling). Their "router" is an oracle reference point where they choose the lower cost model after running both and therefore knowing who passed the test. The cost savings part is only Fireworks theorizing what would happen if an equivale…

> running the task through each model and then picking the cheapest correct option

What system knows what the correct option is and how does it know it?

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#320
post #35

It is not SOTA. Give me a break. Sure, run it on Cerebras to get speed but that’s pretty much its advantage.

Cerebras does not share the quantization of the models so you don't know if you're getting real K3 or k3 lite or something else.

It most likely will be quantized. A cerebras wafer only has 44gb ram, and linking them together vastly reduces the speedup.
Post reply on HN