When DeepSeek was released, it had an immidiate and significant impact on the US stock-market. Now when its becoming common knowledge that China is almost at parity with US SOTA models with good momentum, why is there no sentiment change on the market?
Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
391–400 of 491 posts
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#392If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…
All models are benchmaxxed, period. ”Jagged frontier” is the euphemism du jour, I believe? Anthropic/OpenAI were touting PhD-level intelligence three years ago. And they’re still shipping models that aren’t smart enough to realize things such as the need to drive the car to the car wash (because they hadn’t yet hill-climbed that particular brain-teaser).
No, they weren't. GPT-5 was where OpenAI started talking about PhD-level, and that was less than a year ago.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#393the per-token comparison keeps missing that k3 spends way more tokens per task. if it burns 3x tokens to reach the same result as fable, cheap per-token stops mattering
"So where's this huge price gap coming from? token pricing, prompt caching, and effort-per-task. On SWE for example, K3 works much harder than Fable: roughly 55 turns and 1.3M tokens a task versus 21 turns and 130K. On the long terminal tasks it's the other way around: Fable is the one that spirals, running up 64 turns and 1.5M tokens (sometimes straight into a timeout).
Prompt caching does most of the work of turning that effort into K3's price advantage: even when K3 reads ten times the tokens, with cache hits that means that SWE runs still come in lower cost than Fable. There’s a tradeoff. Tasks with extra turns generally mean more wall-clock time per run i.e. slower runs. If you need an answer in two seconds, that matters; if you're running agents in the background at scale, a bill that's a fraction of the size matters a lot more."
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#394Earlier quoted context omitted.
Could you check how many tokens the different models spent to do the task? With all the comments about Kimi K3 Not being token efficient I'm curious if your test confirms that
For the web app task I mentioned: * Kimi K3: 9532k input (9172k cached), 114k output - cost $5.5 * Qwen 3.8 Max: 18020k input (17836k cached), 114k output - cost $6.3 * Fable: ~14m input (all cached??), 196k output - cost $30 Correction on my earlier post, Kimi was through Pi, not Kimi Code. For Qwen I used Qwen Code and for Fable I used Claude Code. Not sure wtf is going on with the Fable stats (a lot tokens, virtua…
* Kimi K3: (...) cost $6.3
* Fable: (...) cost $30
It's pretty clear that Kimi K3 beats Fable by a long margin.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#395Genuine question: can these posts be paid to hype the open source models? If yes, what would be the purpose? On my work tasks, FastAPI Python and Springboot Java on a modern SaaS product, the only open model that can do tasks well and efficiently is Qwen3.7-Max. In all my experiments, both GLM-5.2 and Kimi are busy grepping around the codebase for ALMOST 70-80K tokens before writing anything and when they do it typic…
Qwen3.7-Max is a proprietary, closed weight model.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#396Earlier quoted context omitted.
The speed at which tokens are crunched, even on the same hardware, differs between models as well. Using more tokens is only a problem if they are processed at the same speed as with a comparison model.
Using more tokens is a significant problem if you pay per token?
Yoh have posts in this thread suggesting that Fable is 5x more expensive than Kimi K3.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#397If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#398When DeepSeek was released, it had an immidiate and significant impact on the US stock-market. Now when its becoming common knowledge that China is almost at parity with US SOTA models with good momentum, why is there no sentiment change on the market?
In reality the world is a highly complex, chaotic, reflexive system, and saying “the entire market moved today because of 12,000,000 factors that randomly aligned” isn’t satisfying enough for people to follow your media channel so they can monetize your eyeballs.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#399Earlier quoted context omitted.
Yes, in the same way I like to kill the enemies in DOOM. It's matrix multiplication. Absurd.
> It's matrix multiplication. Hilarious critique. If you weren't as mathematically illiterate as you likely are, you would know how general matrix operations are, and how essentially any algorithm (including human cognition) can be implemented using them as an intermediate.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#400Earlier quoted context omitted.
I strictly prefer when models ignore any human quirks in my responses. Claude trying to be your friend, saying LOL to your jokes is ridiculous and frankly, harmful
Claude (Opus 4.8) recently told me: > I’d ask you to drop the abuse; (...) if it continues I’ll end the conversation. After I'd used a couple of expletives. And yes it will emit a token. This is truly dystopian. It is NOT a person. What a response. I still cant believe it.