Live data from Hacker News

Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

fireworks.ai

481–490 of 491 posts

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#481
post #469

Earlier quoted context omitted.

Can you elaborate on running them "in conjunction"... are you running the same query on multiple models and then using a third model to judge or make consensus? or am I misunderstanding completely. I'd like to understand how these small models "run together"

I do this in OMP, a fork of Pi. It lets you set different models for different tasks. So with an API that has many different companies models I can set the Plan model to the best one, right now I am using GLM 5.2 for that, it plans really well. I have Vision set to Kimi 2.7 Code (cheaper and vision is just fine). Minimax M3 is set to the Advisor role (double checks work). Deepseek v4 Flash is set for the Task role. A…

Are you using stock OMP or do you have any additional prompts you can share. This looks really promising to me, want to study/learn/copy/steal.... :)

Anything you can share would be useful to me.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#482
post #469

Earlier quoted context omitted.

I do this in OMP, a fork of Pi. It lets you set different models for different tasks. So with an API that has many different companies models I can set the Plan model to the best one, right now I am using GLM 5.2 for that, it plans really well. I have Vision set to Kimi 2.7 Code (cheaper and vision is just fine). Minimax M3 is set to the Advisor role (double checks work). Deepseek v4 Flash is set for the Task role. A…

Are you using stock OMP or do you have any additional prompts you can share. This looks really promising to me, want to study/learn/copy/steal.... :) Anything you can share would be useful to me.

Stock OMP. When you add a provider and then /model you can choose a model and choose what roles it will take. Do that for each role it has.

Then I use /plan when I need to be planning and not writing to any files. Take the advice the UI gives you where it tells you to add the word orchestrate into your prompts when you want to make sure it uses its todo/tasks and sub-agents.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#483

What is routing? How do they decide which is better? The only way to come up with a routing model for your workload is to send queries to all the models and then come up with a way to say which solution was better, often trying out multiple times for the same model + query to account for other statistical errors. This makes you, an ai inference user an unwitting AI company with a non scalable product. The biggest men…

They can but there isn’t evidence that they do.

If they can then that is enough

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#484
post #482

Earlier quoted context omitted.

Are you using stock OMP or do you have any additional prompts you can share. This looks really promising to me, want to study/learn/copy/steal.... :) Anything you can share would be useful to me.

Stock OMP. When you add a provider and then /model you can choose a model and choose what roles it will take. Do that for each role it has. Then I use /plan when I need to be planning and not writing to any files. Take the advice the UI gives you where it tells you to add the word orchestrate into your prompts when you want to make sure it uses its todo/tasks and sub-agents.

I'm off to the races. Thank you.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#485
post #474

Earlier quoted context omitted.

This is such a smooth-brain reply, dude. The capability literally exists, today. You're telling me that it's just SO fantastical that people will be able to catch up? Give me a break. Anthropic and OpenAI are not staffed by demigods.

That's a whole separate argument you've invented here. And if the capability exists today, why did you talk about some nebulous future?

Is it really so hard to understand that open models are behind and yet they are also continually improving and way better than they were a year ago? Like I'm not sure what your point is here. Are you purposely playing dumb to troll people or something?

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#486

If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…

All models are benchmaxxed, period. ”Jagged frontier” is the euphemism du jour, I believe? Anthropic/OpenAI were touting PhD-level intelligence three years ago. And they’re still shipping models that aren’t smart enough to realize things such as the need to drive the car to the car wash (because they hadn’t yet hill-climbed that particular brain-teaser).

> aren’t smart enough to realize things such as the need to drive the car to the car wash

When this first went viral, I immediately tried it on Opus (whatever version was latest at that time), and it got it first try. Tried a few more times in fresh sessions and it got it every time.

Sonnet did screw it up, though.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#487
post #397

If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…

There are some benchmarks I personally trust. But the most reliable indicator other than public benchmarks is the quality of products people are working on, and the model they use for the work. Not a benchmark scoreboard or a one~few shot demo, but the product they have been building for weeks. The reality is, even the diehard advocates who build and sell tooling for open weights, still use proprietary frontier model…

Sure, the tools and models I use to build real products says a lot about how much I trust them. The closed frontier models are the most familiar, but I'm trying to unlearn my dependence on them. I just don't like them, and I believe AI should be open. In the short term I may lose out on some coding productivity, but in the long term I want to be a better AI engineer anyway. Open model + configurable harness forces me to learn how to get things done more efficiently, and understand how harnesses and models work under the hood.

I'll see what I can ship while not relying on these closed frontier models at all, as unreasonable as that might be. And if I don't get any customers that's on me.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#488
post #400

Earlier quoted context omitted.

Claude (Opus 4.8) recently told me: > I’d ask you to drop the abuse; (...) if it continues I’ll end the conversation. After I'd used a couple of expletives. And yes it will emit a token. This is truly dystopian. It is NOT a person. What a response. I still cant believe it.

You made the model cry :(

I don't berate or insult the models. If you treat your computer like this you might get into habit of it with people as well.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#490
post #459

Earlier quoted context omitted.

The comment I replied to didn't mention anything about being close to frontier, just a blanket statement about Chinese Labs models being slow and inefficient. People talk about frontier as if it's the only innovation worth pursuing. Deepseek V4 is far from fontier, but it's architecture is super innovative and efficient and what it achieves at that size (especially V4 Flash) is incredible.

> Having tested K3, Qwen 3.8 max preview, Fable and Sol for the past few days, the Chinese models really are slow and token-inefficient. Those are all frontier-competitive models.

I guess Deepseek V4 is too now that it's out of preview.
Post reply on HN