Live data from Hacker News

Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

fireworks.ai

431–440 of 491 posts

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#431
post #276

If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…

Having tested K3, Qwen 3.8 max preview, Fable and Sol for the past few days, I partially agree. Don’t trust the benchmarks, and the Chinese models really are slow and token-inefficient. However they do seem very close to SOTA : I’d say roughly equal to previous gen (Opus 4.8, GPT 5.5). It’s yet another silly benchmark, but compare them here: https://senko.net/vibecode-bench/ I also had K3, Qwen3.8 and Fable (using Ki…

> the Chinese models really are slow and token-inefficient

This totally depends on the model. Deepseek V4 is very fast and efficient.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#433
post #344

This benchmark is probably also self promotion of services. Fireworks happens to make a router. Using the router gets better performance. https://docs.fireworks.ai/deployments/routers

I work for a router company too, I ran some tests on all of the cheapest models and came to the same outcome where a handful of small models ran together in conjunction outperform SoTA models -- outperforms in that it got a 95% vs a 94% and I bet that changes with the day of the week. Anyways, I did get a similar result in a different sort of measurement.

Can you elaborate on running them "in conjunction"... are you running the same query on multiple models and then using a third model to judge or make consensus? or am I misunderstanding completely. I'd like to understand how these small models "run together"

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#434

When DeepSeek was released, it had an immidiate and significant impact on the US stock-market. Now when its becoming common knowledge that China is almost at parity with US SOTA models with good momentum, why is there no sentiment change on the market?

deepseek drop was based on assumption that we overestimated how much compute and infra is needed to train models. They seem to have claimed that they did it in couple of million or something.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#435

Earlier quoted context omitted.

Why? LLMs are not humans.

Doors aren't humans either, yet we design their handles and locks to be graspable and manipulable by humans. The purpose of technology is to serve humans. Therefore, technology must conform as much as possible to human sensibilities rather than vice versa.

Yet handles are not shaped like hands. The speakers I use on my computer don't look anything like a mouth, nor does the microphone I use on calls look like a set of ears.

Ultimately I guess this is up to opinion, but in mine, humanizing LLMs is not exactly "conforming to human sensibilities", rather it's trying to pass the LLM as something else to make it more appealing. It is deceitful in this way, and that I completely abhor.

For example, "Please" and "Thank you" come from human sentiment. It's an expression of something underneath, and an LLM using such expressions not only is fake, but makes a mockery of the real thing.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#436
post #413

Earlier quoted context omitted.

Anthropic/OpenAI were touting PhD-level intelligence three years ago No, they weren't. GPT-5 was where OpenAI started talking about PhD-level, and that was less than a year ago.

(March 2024) Antropic claimed Claude 3 Opus had "graduate-level expert reasoning" with GPQA results of around 60% showing a roughly phd level performance. (Sept 2024) OpenAI claimed o1 was phd-level in their launch post. You're kinda both wrong. :)

They claimed the model was PhD-level, but they never mentioned the university the model graduated from... :)

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#437
post #431
post #276

Earlier quoted context omitted.

Having tested K3, Qwen 3.8 max preview, Fable and Sol for the past few days, I partially agree. Don’t trust the benchmarks, and the Chinese models really are slow and token-inefficient. However they do seem very close to SOTA : I’d say roughly equal to previous gen (Opus 4.8, GPT 5.5). It’s yet another silly benchmark, but compare them here: https://senko.net/vibecode-bench/ I also had K3, Qwen3.8 and Fable (using Ki…

> the Chinese models really are slow and token-inefficient This totally depends on the model. Deepseek V4 is very fast and efficient.

Not close to frontier.

But yes, it’s cheap.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#438

When DeepSeek was released, it had an immidiate and significant impact on the US stock-market. Now when its becoming common knowledge that China is almost at parity with US SOTA models with good momentum, why is there no sentiment change on the market?

The “significant impact” lasted like 2 days and then the market corrected itself.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#439
Seems that “oracle routing” is a term the authors invented. Also sounds like they’re sending requests to all routes but assuming the readers request will route to the “best” model. The closest thing to what they’re describing is semantic routing using a NN search in a vector DB to make a routing decision, but the efficacy of this approach isn’t a slam dunk.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#440
post #251

Earlier quoted context omitted.

Is that any worse than being subsidized by VC? They're playing the same game as the western labs, just with slightly different players.

Chinese government vs American VCs doesn't equate to "slight different players."

American VCs and American Government mingle at the deepest levels. Yes, slightly different is correct.
Post reply on HN