Live data from Hacker News

Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

fireworks.ai

381–390 of 491 posts

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#381

Earlier quoted context omitted.

It doesn’t truly feel anything. It will adopt whatever tone it’s prompted to.

Well yes. But some of the glazing comes from system prompts/training they do before end users get their hands on prompting it. Of course, you can try to make it ruder than vanilla if you wish (I recommend). The question is if the producers of these models were less incentivized to make them agreeable simply because most people don't like being spoken to like an idiot (or having their asks vetoed), how would they actu…

I think it is a huge mistake to believe that there is a “true” nature of the models. They are artificial, everything about them is a construct. This is like asking the true shape of the clay without the potter’s influence.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#384

we put a router model in front of two other models so the router can decide which model is better at deciding things. next we'll need a router for the router and eventually the entire internet is just routers routing routers to other routers

It's all pipes, Jerry.

Jerry, these are load bearing routers.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#385

If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…

I've seen this phrase repeated over and over again and I don't think it is true, or at least, to the degree that "benchmaxxing" claims are true, American frontier AI companies are probably about as guilty of it as Chinese AI companies are.

For both GLM 5.2 and Kimi K3 I feel the rough average of the benchmarks gives you a rough idea where they stand. GLM 5.2 was sitting somewhere behind Opus 4.8 but it didn't feel very far away. I've used Kimi K3 via OpenRouter and while I've had limited experience so far, it sure as shit feels like it's right up there with Sol and Fable to me. I am happily able to believe Fable has the edge still, but on a request by request basis it would be pretty easy to get an impression one way or another.

The existing Kimi models were already pretty good so I really just don't find this new model to be that hard to believe. Maybe I'm naive.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#386
post #276

If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…

Having tested K3, Qwen 3.8 max preview, Fable and Sol for the past few days, I partially agree. Don’t trust the benchmarks, and the Chinese models really are slow and token-inefficient. However they do seem very close to SOTA : I’d say roughly equal to previous gen (Opus 4.8, GPT 5.5). It’s yet another silly benchmark, but compare them here: https://senko.net/vibecode-bench/ I also had K3, Qwen3.8 and Fable (using Ki…

I love Kimi for some things, but K3 to me struggles with things Opus 4.8 breezes through. Admittedly stuff that is on the more complex side (code generation bugs in a compiler) to the point that after two days of struggle I stopped Kimi and will have Opus redo its work once I have spare tokens...

This wasn't just being slow - it didn't make forward progress.

For simpler stuff even 2.7 does just fine, though.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#389

When DeepSeek was released, it had an immidiate and significant impact on the US stock-market. Now when its becoming common knowledge that China is almost at parity with US SOTA models with good momentum, why is there no sentiment change on the market?

Because, as you said, this is no longer news. Deepseek was the news. This is just the predictable progress playing out.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#390

Earlier quoted context omitted.

It doesn’t truly feel anything. It will adopt whatever tone it’s prompted to.

Well yes. But some of the glazing comes from system prompts/training they do before end users get their hands on prompting it. Of course, you can try to make it ruder than vanilla if you wish (I recommend). The question is if the producers of these models were less incentivized to make them agreeable simply because most people don't like being spoken to like an idiot (or having their asks vetoed), how would they actu…

While I wish the models were glazing less, I don't think I want them pushing back more until they get smarter. The current Opus situation where it questions good, thought-out decisions is a bit mad. It may be better if the model asked about your thought process or motivations behind certain things but I don't think I want to explain myself to an LLM all the time either.

Recent example: I asked Opus for some kind of reference data for MTBF in software. It took a few minutes to do "research", ended up providing zero data, and gave me a long essay on how DORA is superior and I should be using that instead. When I clarified that we're thinking of adding MTBF next to our existing DORA metrics, it decided I need to be thought about SLOs... if it was the first time, I may have continued clarifying but since it wasn't, I just gave up on Claude for this topic.

Post reply on HN