Earlier quoted context omitted.
It doesn’t truly feel anything. It will adopt whatever tone it’s prompted to.
Well yes. But some of the glazing comes from system prompts/training they do before end users get their hands on prompting it. Of course, you can try to make it ruder than vanilla if you wish (I recommend). The question is if the producers of these models were less incentivized to make them agreeable simply because most people don't like being spoken to like an idiot (or having their asks vetoed), how would they actu…
Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
381–390 of 491 posts
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#382Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#383Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#384we put a router model in front of two other models so the router can decide which model is better at deciding things. next we'll need a router for the router and eventually the entire internet is just routers routing routers to other routers
It's all pipes, Jerry.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#385If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…
For both GLM 5.2 and Kimi K3 I feel the rough average of the benchmarks gives you a rough idea where they stand. GLM 5.2 was sitting somewhere behind Opus 4.8 but it didn't feel very far away. I've used Kimi K3 via OpenRouter and while I've had limited experience so far, it sure as shit feels like it's right up there with Sol and Fable to me. I am happily able to believe Fable has the edge still, but on a request by request basis it would be pretty easy to get an impression one way or another.
The existing Kimi models were already pretty good so I really just don't find this new model to be that hard to believe. Maybe I'm naive.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#386If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…
Having tested K3, Qwen 3.8 max preview, Fable and Sol for the past few days, I partially agree. Don’t trust the benchmarks, and the Chinese models really are slow and token-inefficient. However they do seem very close to SOTA : I’d say roughly equal to previous gen (Opus 4.8, GPT 5.5). It’s yet another silly benchmark, but compare them here: https://senko.net/vibecode-bench/ I also had K3, Qwen3.8 and Fable (using Ki…
This wasn't just being slow - it didn't make forward progress.
For simpler stuff even 2.7 does just fine, though.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#387Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#388Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#389When DeepSeek was released, it had an immidiate and significant impact on the US stock-market. Now when its becoming common knowledge that China is almost at parity with US SOTA models with good momentum, why is there no sentiment change on the market?
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#390Earlier quoted context omitted.
It doesn’t truly feel anything. It will adopt whatever tone it’s prompted to.
Well yes. But some of the glazing comes from system prompts/training they do before end users get their hands on prompting it. Of course, you can try to make it ruder than vanilla if you wish (I recommend). The question is if the producers of these models were less incentivized to make them agreeable simply because most people don't like being spoken to like an idiot (or having their asks vetoed), how would they actu…
Recent example: I asked Opus for some kind of reference data for MTBF in software. It took a few minutes to do "research", ended up providing zero data, and gave me a long essay on how DORA is superior and I should be using that instead. When I clarified that we're thinking of adding MTBF next to our existing DORA metrics, it decided I need to be thought about SLOs... if it was the first time, I may have continued clarifying but since it wasn't, I just gave up on Claude for this topic.