Why SoTA (uppercase “T”) instead of SotA (lowercase “T”) ? “State of [T]he Art” versus “State of [t]he Art”. If not SotA then at least SOTA, which is more accurate.
Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
71–80 of 491 posts
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#72A third the cost, open source, and won't refuse every other request because of some vague possible connection to cybersecurity concerns.
From the blog post moonshot refers to it as open source but only mentions releasing the weights.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#73Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#74[flagged]
Moonshot has committed to release it end of the month.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#75I love the Chinese models. I use DeepSeek exclusively and now Kimi K3 offers a great planning assistant for more advanced coding tasks. DeepSeek v4 Flash is extremely fast and is able to handle pretty much anything I've thrown at it (I use mostly Rust, PSQL, Angular and Terraform). I self host Bifrost as my LLM gateway, though I wish LLM vendors would do monthly/daily automatic billing (like VPS providers do) rather…
The best part is using harnesses like reasonix or whale make cache hit at a rate close to 98%, making requests converge to practically free. And that's with unsubsidized American providers like cloudflare or Digital Ocean.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#76A third the cost, open source, and won't refuse every other request because of some vague possible connection to cybersecurity concerns.
My questions about strawberries got blocked as too dangerous! I’m not joking
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#77Earlier quoted context omitted.
The weights are the "source" of a model.
The weights are the output artifact, the training corpus and system are the source.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#78Earlier quoted context omitted.
Strawberries are a well known weakness of LLMs, as they have a hard time to count the numbers of "r"s in them. Maybe that's why, because they fear that weakness could be exploited somehow.
Probably need to be taught by someone of Latino origin. Learning to roll them "r"s could help.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#79It is not SOTA. Give me a break. Sure, run it on Cerebras to get speed but that’s pretty much its advantage.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#80Very interesting. They test Kimi K3 and Fable on a set of approx 1000 tasks grouped into 5 areas (SWE, Legal, etc). They put a router model in front that predicts whether Kimi or Fable is going to give a better cost for a correct result. (They believe that ultimately such a router model should be continuously trained on your own workloads so it makes the best decisions for you). Their router chose Kimi the majority o…
Their "router" is an oracle reference point where they choose the lower cost model after running both and therefore knowing who passed the test. The cost savings part is only Fireworks theorizing what would happen if an equivalent predicting router exists. That's a big if.