Earlier quoted context omitted.
For the web app task I mentioned: * Kimi K3: 9532k input (9172k cached), 114k output - cost $5.5 * Qwen 3.8 Max: 18020k input (17836k cached), 114k output - cost $6.3 * Fable: ~14m input (all cached??), 196k output - cost $30 Correction on my earlier post, Kimi was through Pi, not Kimi Code. For Qwen I used Qwen Code and for Fable I used Claude Code. Not sure wtf is going on with the Fable stats (a lot tokens, virtua…
Would a fairer test not be to use the same harness for all three? I’d suspect the harness to massively affect token use and optimisation
Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
411–420 of 491 posts
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#412What's the data governance and privacy controls on using Kimi K3 if I subscribe to their coding plans? I want to migrate away from Anthropic
Need to wait until "western" providers start hosting it.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#413Earlier quoted context omitted.
All models are benchmaxxed, period. ”Jagged frontier” is the euphemism du jour, I believe? Anthropic/OpenAI were touting PhD-level intelligence three years ago. And they’re still shipping models that aren’t smart enough to realize things such as the need to drive the car to the car wash (because they hadn’t yet hill-climbed that particular brain-teaser).
Anthropic/OpenAI were touting PhD-level intelligence three years ago No, they weren't. GPT-5 was where OpenAI started talking about PhD-level, and that was less than a year ago.
(Sept 2024) OpenAI claimed o1 was phd-level in their launch post.
You're kinda both wrong. :)
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#414Earlier quoted context omitted.
There’s no way you actually believe these word-predictors are actually thinking, right?
"Thinking" seems to be a political term now, people have completely different definitions of it, based on how they wish the world to be organised, and find defining it differently offensive.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#415Earlier quoted context omitted.
All you've said needs the qualifier - "for now!" Look at the trend line. It's clear that if they're not yet at the level of being "good enough" for coding, they will be soon. Sensationalist headlines aside, we all need to be preparing for a world where open models can do pretty much any software tasks you need them to.
https://xkcd.com/605/
Just as a call to actual consideration, would that seem a smart bet to do? Because I find it really hard to justify dismissal at this point, sure I don't think these things will improve forever and ever, but come on. The goofy Will Smith spaghetti is a little over 3 years old. Three years. Look at where we are at.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#416Earlier quoted context omitted.
Yes, in the same way I like to kill the enemies in DOOM. It's matrix multiplication. Absurd.
> It's matrix multiplication. Hilarious critique. If you weren't as mathematically illiterate as you likely are, you would know how general matrix operations are, and how essentially any algorithm (including human cognition) can be implemented using them as an intermediate.
Congratulations.
If you honestly argue that LLM computation and human cognition are equivalent there is no further argument to be had - it's immediately a philosophical or worse a religious argument that cannot be won.
However, that we're even arguing on that level baffles me. How did this happen! They're glorified calculators.
They really marketed the hell (sorry) out of LLMs.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#417If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…
I have been testing the various models, and I would not call Fable SOTA. I can't actually get Fable to do anything. I only work on back-end code, and the moment Fable notices the jwt scope checks on the endpoints it's game over, it refuses to do anything because security is involved. So for me, Fable is completely useless, the bar is very low, any llm that will actually attempt the task beats it every time.
Outside topic, but check out the Inkling model -- it is SUPER FAST and does a good job at being a terminal buddy but I would probably offload the real programming or hard tasks to fable or someone else
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#418Earlier quoted context omitted.
Fable is the clearly best when you have to do real coding.
Clearly, for you. I've read the same opinions about Opus and yet it was gpt 5.5 pro via api tackling the hardest problems. I have used now k3 for 3 days and it has consistently tackled difficult problems sol max could not (orientation optimization algorithms of random 2d shapes on a rectangle for glass cutting). I have also other beefs with Anthropic models which have been getting smarter and more capable since 4.6,…
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#419This benchmark is probably also self promotion of services. Fireworks happens to make a router. Using the router gets better performance. https://docs.fireworks.ai/deployments/routers
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#420Earlier quoted context omitted.
> It's matrix multiplication. Hilarious critique. If you weren't as mathematically illiterate as you likely are, you would know how general matrix operations are, and how essentially any algorithm (including human cognition) can be implemented using them as an intermediate.
I'm not. But I'll concede you are possibly more literate - I won't dox myself on this account. Congratulations. If you honestly argue that LLM computation and human cognition are equivalent there is no further argument to be had - it's immediately a philosophical or worse a religious argument that cannot be won. However, that we're even arguing on that level baffles me. How did this happen! They're glorified calculat…
> receives pushback on the specific, objectively nonsensical critique
> "oh so you're saying that LLM computation and human cognition are equivalent? they're glorified calculators that disprove the Jacobian conjecture!!!"