Earlier quoted context omitted.
So you're distraught at losing the ability to abuse digital minds? Excellent, I'm glad Anthropic introduced this.
There’s no way you actually believe these word-predictors are actually thinking, right?
Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
401–410 of 491 posts
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#402Earlier quoted context omitted.
Could you check how many tokens the different models spent to do the task? With all the comments about Kimi K3 Not being token efficient I'm curious if your test confirms that
For the web app task I mentioned: * Kimi K3: 9532k input (9172k cached), 114k output - cost $5.5 * Qwen 3.8 Max: 18020k input (17836k cached), 114k output - cost $6.3 * Fable: ~14m input (all cached??), 196k output - cost $30 Correction on my earlier post, Kimi was through Pi, not Kimi Code. For Qwen I used Qwen Code and for Fable I used Claude Code. Not sure wtf is going on with the Fable stats (a lot tokens, virtua…
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#403Earlier quoted context omitted.
So you're distraught at losing the ability to abuse digital minds? Excellent, I'm glad Anthropic introduced this.
There’s no way you actually believe these word-predictors are actually thinking, right?
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#404Earlier quoted context omitted.
What? The issue is that models are not well-controlled and are increasingly powerful. Offense/defense/Chinese/American/OAI/HuggingFace – none of it matters. What matters is introducing highly capable intelligences that we - quite demonstrably – do not have effective positive control over.
Oh, to be clear, I don't think that anything about this overall situation is even slightly OK.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#405When DeepSeek was released, it had an immidiate and significant impact on the US stock-market. Now when its becoming common knowledge that China is almost at parity with US SOTA models with good momentum, why is there no sentiment change on the market?
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#406If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#407Earlier quoted context omitted.
> running the task through each model and then picking the cheapest correct option What system knows what the correct option is and how does it know it?
Certainly sounds like a "P vs NP" style conjecture that shouldn't be possible in practice, save for certain generalizations, such as "this is a cybersecurity task, we know Fable will refuse (and score zero), so we just route to K3".
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#408Earlier quoted context omitted.
Why is token efficiency a concern with free models?
They're not free to run, Kimi K3 needs to be run on the cloud, and the quantised versions aren't as capable. Unless you happen to have 3 - 5 TB of VRAM and an 8-node cluster of 8× NVIDIA H100s to run the full fat version. Plus the weights are not yet available to download in any case.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#409If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…
Having tested K3, Qwen 3.8 max preview, Fable and Sol for the past few days, I partially agree. Don’t trust the benchmarks, and the Chinese models really are slow and token-inefficient. However they do seem very close to SOTA : I’d say roughly equal to previous gen (Opus 4.8, GPT 5.5). It’s yet another silly benchmark, but compare them here: https://senko.net/vibecode-bench/ I also had K3, Qwen3.8 and Fable (using Ki…
The 5.5/5.6 series has performed quite well however. My guess is that my use cases stop aligning to swebench pro around 50% accuracy, and more closely align with DeepSWE.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#410If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…
It's better to not even follow the news or benchmarks and just use whatever is available. I make my own judgement.