Live data from Hacker News

Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

fireworks.ai

81–90 of 491 posts

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#81
post #41

Earlier quoted context omitted.

From the blog post moonshot refers to it as open source but only mentions releasing the weights.

The weights are the "source" of a model.

I want to caution against this line of thinking. NSA's fast16 program silently altered data during nuclear simulations, and that is also possible within open weight models. There is no reason to think there is anything like that currently, it also cannot be dismissed. Open source would include the training data so you could create a similarly capable model.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#82

I love the Chinese models. I use DeepSeek exclusively and now Kimi K3 offers a great planning assistant for more advanced coding tasks. DeepSeek v4 Flash is extremely fast and is able to handle pretty much anything I've thrown at it (I use mostly Rust, PSQL, Angular and Terraform). I self host Bifrost as my LLM gateway, though I wish LLM vendors would do monthly/daily automatic billing (like VPS providers do) rather…

Prepaid means someone can't sign up, use a bunch of inference, then cancel their card and disappear into the sunset. VPS is a more long term investment where it's harder to switch and it doesn't cost the provider much if a few users jump out without paying for a month.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#83
post #42

I love the Chinese models. I use DeepSeek exclusively and now Kimi K3 offers a great planning assistant for more advanced coding tasks. DeepSeek v4 Flash is extremely fast and is able to handle pretty much anything I've thrown at it (I use mostly Rust, PSQL, Angular and Terraform). I self host Bifrost as my LLM gateway, though I wish LLM vendors would do monthly/daily automatic billing (like VPS providers do) rather…

What don't you like about the service, besides the mark up?

For me, the only utility OpenRouter gives me is billing consolidation - I don't really need the routing capabilities because I use Bifrost for that.

As a router, it's not very feature rich. For example I restricted the available models to the ones I want to use however the `/models` endpoint still lists all the models, making my LLM client list the 200+ models available on the service (even though they will throw an error if I try to use them).

With Bifrost, I can also create model aliases with custom configuration - for example I can create a model alias `deepseek-v4-flash-nothink` which disables thinking. I can create `deepseek-v4-flash-caveman` which injects the caveman skill (to save tokens) etc.

Plus I can contribute to Bifrost, which I can't do with OpenRouter.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#84

Earlier quoted context omitted.

The weights are the "source" of a model.

If weights are the source for models then ELF binaries are the source for software.

clearly not true. the weights are the preferred form for making modifications. Do you really think people should be downloading hundreds of TB of training data and running make to build the model on their own cluster of GPUs?

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#85
post #33

Earlier quoted context omitted.

$20 users don't get access to Fable. It's $100+ tier only.

True. However, I believe most non-programmers don't need access to the fanciest model, but just want to use a good LLM without constant nagging about usage limits. Then 19 vs. 20 is true?

No "19 vs 20" is not true for "Kimi vs Fable". It might be true for "Kimi vs Opus" or whatever. Anyways looking at pricing plans is not a good way of comparing the price of LLMs. It makes more sense to look at cost per token:

  Kimi K3: $3/$15 (input/output)
  Fable:  $10/$50

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#86

Earlier quoted context omitted.

Interestingly the reality of the open source release is they are opening up to full distillation by the closed source model providers at a deeper and more fundamental level. If anything the open sourcing will help Anthropic and open ai ladder up faster. Open source has always been about mutual cooperation towards a goal and has never closed the door to commercial success. All the hand wringing about open weight model…

The closed labs don't really benefit unless the open model has something extra they don't have though. Meanwhile the open model dilutes their customer base and seriously cheapens their offering (which is a heck of a good though IMO).

Because the open model might have been distilled from outputs of the closed model doesn’t mean they are architecturally equivalent. There is almost certainly innovations in architecture present in the open models that the closed labs didn’t think of. It also is almost certainly true that they aren’t completely built out of a distilled corpus, that reinforcement is equivalent, etc. Therefore closed labs will also benefit from being able to inspect in totality the architecture, activations, weights, and be able to train against it at scale in an ensemble of other models and their internal work.

The only way there is no benefit would be is if the open models are literal copies of the closed model, which unless there was direct theft, is highly improbable.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#87
post #62

Earlier quoted context omitted.

The weights are the "source" of a model.

The weights are the output artifact, the training corpus and system are the source.

The training corpus + training code is just an automated editor for the weights. It would be like requiring the source code for Visual Studio for software made within it to be open source.

Weights are not an output artifact no more than source code is. There is never a moment where you can claim that it's done. As requirements change new ways to change the weights / code come up. With different projects you might want to import the weights / code into a bigger model / codebase.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#88
post #10

Hmmm, a company that hosts open models is telling us how good open models are...

They share their methodology and results. I learned things about the relative strengths and weaknesses of Kimi and Fable I hadn’t seen anywhere else. Should being in the model hosting business disqualify them from sharing?

Doesn't disqualify them, but it may call into question their results seeing as they have a potential conflict of interest.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#89
post #10

Hmmm, a company that hosts open models is telling us how good open models are...

The don't only host open weight models. Also, why not promote this. If Fireworks thinks this big news might convert some new business doesn't make it not true.

I'd be doing the same thing if I were them as a marketing move

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#90

Earlier quoted context omitted.

Interestingly the reality of the open source release is they are opening up to full distillation by the closed source model providers at a deeper and more fundamental level. If anything the open sourcing will help Anthropic and open ai ladder up faster. Open source has always been about mutual cooperation towards a goal and has never closed the door to commercial success. All the hand wringing about open weight model…

I'm happy that Kimi K3 is indeed SotA and its open weights are due to be released soon. It's also true that Moonshot and other labs distill from Claude. This has been reported on extensively. I don't think there's any alpha for Anthropic distilling from this model. I do not mean to discount the tremendous amount of innovation regarding MoE and quantization that Moonshot has accomplished. But its training with synthet…

When I say distill I also mean mine it architecturally for insights but I doubt seriously the model training is entirely distillation of Claude, it’s almost certainly a mixture of both original corpus and reinforcement as well as distillation. I think it’s a little condescending to imply that these new open models are cheap ripoffs with nothing original to them. These teams and labs are top tier as well, working under unreasonable constraints imposed by the USG. That’s a powerful combination for creativity.
Post reply on HN