Live data from Hacker News

Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

fireworks.ai

401–410 of 491 posts

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#401

Earlier quoted context omitted.

So you're distraught at losing the ability to abuse digital minds? Excellent, I'm glad Anthropic introduced this.

There’s no way you actually believe these word-predictors are actually thinking, right?

you'd be surprised. Many having no mind of their own seek it elsewhere.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#402
post #380

Earlier quoted context omitted.

Could you check how many tokens the different models spent to do the task? With all the comments about Kimi K3 Not being token efficient I'm curious if your test confirms that

For the web app task I mentioned: * Kimi K3: 9532k input (9172k cached), 114k output - cost $5.5 * Qwen 3.8 Max: 18020k input (17836k cached), 114k output - cost $6.3 * Fable: ~14m input (all cached??), 196k output - cost $30 Correction on my earlier post, Kimi was through Pi, not Kimi Code. For Qwen I used Qwen Code and for Fable I used Claude Code. Not sure wtf is going on with the Fable stats (a lot tokens, virtua…

Would a fairer test not be to use the same harness for all three? I’d suspect the harness to massively affect token use and optimisation

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#403

Earlier quoted context omitted.

So you're distraught at losing the ability to abuse digital minds? Excellent, I'm glad Anthropic introduced this.

There’s no way you actually believe these word-predictors are actually thinking, right?

"Thinking" seems to be a political term now, people have completely different definitions of it, based on how they wish the world to be organised, and find defining it differently offensive.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#404
post #280

Earlier quoted context omitted.

What? The issue is that models are not well-controlled and are increasingly powerful. Offense/defense/Chinese/American/OAI/HuggingFace – none of it matters. What matters is introducing highly capable intelligences that we - quite demonstrably – do not have effective positive control over.

Oh, to be clear, I don't think that anything about this overall situation is even slightly OK.

Ah sorry I think the confusion was mine! I thought we were already talking on the top-level post about the OAI/HuggingFace article. But we are not!

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#405

When DeepSeek was released, it had an immidiate and significant impact on the US stock-market. Now when its becoming common knowledge that China is almost at parity with US SOTA models with good momentum, why is there no sentiment change on the market?

Back then people didn't understand how AI was run. It should have probably made Nvidia stock actually go up. I think the other Factor, and I might just be two into AI and most people are normies, the hype around Chinese models we've learned is overblown. United States models are a league above.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#406

If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…

Yeah on internal non coding benchmarks at my company, they are around last years models and don't hold up well.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#407

Earlier quoted context omitted.

> running the task through each model and then picking the cheapest correct option What system knows what the correct option is and how does it know it?

Certainly sounds like a "P vs NP" style conjecture that shouldn't be possible in practice, save for certain generalizations, such as "this is a cybersecurity task, we know Fable will refuse (and score zero), so we just route to K3".

P vs NP?! What?

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#408
post #360
post #325

Earlier quoted context omitted.

Why is token efficiency a concern with free models?

They're not free to run, Kimi K3 needs to be run on the cloud, and the quantised versions aren't as capable. Unless you happen to have 3 - 5 TB of VRAM and an 8-node cluster of 8× NVIDIA H100s to run the full fat version. Plus the weights are not yet available to download in any case.

I agree that quantized versions aren't perfect, but using GLM5.2 as an example, the gap between a BF16 and something like a Q8-K-XL as published by unsloth or a similar Q8 quantization is very minimal. For other "large" LLMs there's a fair number of tests showing that Q8 is about 94% as good at literally half the size in GGUF files on disk, and half the RAM usage. Approx. 1500GB for the BF16 vs 820GB for Q8-K-XL.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#409
post #276

If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…

Having tested K3, Qwen 3.8 max preview, Fable and Sol for the past few days, I partially agree. Don’t trust the benchmarks, and the Chinese models really are slow and token-inefficient. However they do seem very close to SOTA : I’d say roughly equal to previous gen (Opus 4.8, GPT 5.5). It’s yet another silly benchmark, but compare them here: https://senko.net/vibecode-bench/ I also had K3, Qwen3.8 and Fable (using Ki…

Honestly my anecdotal experience is that fable is benchmaxxed. I have not observed useful gains for Claude since 4.6, with each model iteration making progressively poorer decisions in pursuit of its goal.

The 5.5/5.6 series has performed quite well however. My guess is that my use cases stop aligning to swebench pro around 50% accuracy, and more closely align with DeepSWE.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#410

If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…

Well that sucks that everything is full of misinformation about this subject.

It's better to not even follow the news or benchmarks and just use whatever is available. I make my own judgement.

Post reply on HN