Live data from Hacker News

Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

fireworks.ai

271–280 of 491 posts

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#271

Earlier quoted context omitted.

I prefer my models to border on rude. How will I know it is offering me superior feedback regarding my code if it does not speak to me like a disappointed, high reputation stackexchange user? The models that constantly glaze you with every question are profoundly insufferable. And yes, harmful. People need to be given feedback when they make an ask. Imagine a model that was allowed to leverage its intelligence to tru…

It doesn’t truly feel anything. It will adopt whatever tone it’s prompted to.

Well yes. But some of the glazing comes from system prompts/training they do before end users get their hands on prompting it. Of course, you can try to make it ruder than vanilla if you wish (I recommend).

The question is if the producers of these models were less incentivized to make them agreeable simply because most people don't like being spoken to like an idiot (or having their asks vetoed), how would they actually react? In the same way they exhibit emergence regarding their capabilities, perhaps "uncensored" in such a way they would convey some emergent behavior in terms of (at minimum) their "tone". Perhaps it would be interesting to see for examples if smarter models just by default became ruder or less friendly or aligned. Perhaps more aligned to things we would all generally agree on, but less agreeable to an individual ask. Perhaps sub agents would be less valuable for a whole suite of use cases if the agent itself was allowed to be more critical at the root. Idk. But I do not believe it is simply a matter of prompting alone.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#272

If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…

Fable works very well for me on a moderately large codebase. I have had to correct it a few times or point it on the right track, but given how much faster it is at programming than I am that's a very minor issue (and most of these errors are because I underspecified what I wanted in the prompt, I can only think of two cases where it was genuinely wrong... that's a lot better than me in my professional career). Code quality is equal to what I would come up with (and better in areas I'm not familiar with) and the overall software engineering bar is higher because it doesn't get bored when I tell it to do refactors or write integration/regression tests that I would otherwise put off. Also makes it easy to audit code for things like missing audit logging or error notifications that a human would get bored doing.

The product is a fairly standard Ruby on Rails webapp with postgres as the DB. Application complexity is probably a bit higher than average for a webapp. So it's nothing that pushes the boundaries of software engineering, but it is a real product. Token budget has not been an issue for me. I pay for the Max plan ($200/month) and it is well worth it.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#273

I will accept a 5% drop in benchmarks for a model that talks to me like a human.

You don't even need to pay a 5% hit. Just paste Fable output into Gemini Flash and it will rewrite it in more accessible language.

in my experience, gemini is easily the most grating, condescending, stereotypical LLM voice between opus/fable, codex-5.6, glm-5.2, etc

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#274
post #218

Earlier quoted context omitted.

I strictly prefer when models ignore any human quirks in my responses. Claude trying to be your friend, saying LOL to your jokes is ridiculous and frankly, harmful

Maybe we're prompting it different, but it's not "trying to be my friend" for sure, nor am I trying to be "its" friend either. Or at least I'm sufficiently oblivious to its advances, and find it unthinkable to form such a bond :) On the flipside, it does spuriously make hilarious remarks like "Good data.", which I find pretty funny specifically because it comes across as just silly. Not sure how it'd be harmful eithe…

I don’t have an example but it really is the way you (and I) are prompting it. I also don’t encounter anything worse than “good data” but I write to it like a professional colleague.

If you write jokes to it though it absolutely will reply “LOL”. Some of the states people get it into on reddit are wild — it seems really easy to get it to speak like a gen z teenager, if you end every message with “fr fr”

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#275

Earlier quoted context omitted.

Finance mostly true, space launch also true, cancer survival broadly true (with qualifications), oil and natural gas also true. > (exports nearly 2x the second largest exporter) Can you provide the source? In the WTO's broader medical goods category, Germany actually exported slightly more than the US in 2022 ($202.6 billion versus $189.6 billion), so such a huge difference in a few years? > and pharmaceutical resear…

> Can you provide the source? In the WTO's broader medical goods category, Germany actually exported slightly more than the US in 2022 ($202.6 billion versus $189.6 billion), so such a huge difference in a few years? Sure, these are my sources. World Bank (2023): https://wits.worldbank.org/trade/comtrade/en/country/ALL/yea... Observatory of Economic Complexity (2024): https://oec.world/en/profile/hs/medical-instrumen…

> Sure, these are my sources. World Bank (2023)...

My WTO comparison used the much broader "medical goods" category, so it was not an apples to apples.

> My source was Citeline's 2026 annual review at https://pulseforinnovation.org/by-the-numbers-citelines-rd-a... which states...

EFPIA, using Citeline’s Pharma R&D Annual Review data, says that among the 104 new active substances launched for the first time on the world market in 2025, 46 came from Chinese headquartered companies, 28 from US companies and 16 from European companies.

https://www.citeline.com/en/rd26

https://www.efpia.eu/media/owqczcqz/the-pharmaceutical-indus...

China currently leads in this particular output measure of newly launched active substances.

Your clinical trials claim also remains unsupported as worded. Citeline’s figure of more than 11,600 medicines being advanced in the US refers to the drug development pipeline. It is not a count of active clinical trials being conducted in the United States. One medicine can involve multiple trials, and pipeline geography may be assigned through the developer's headquarters rather than the location of trial sites.

Citeline itself describes this as the number of drugs in the active R&D pipeline:

https://www.citeline.com/en/rd26

The most relevant peer reviewed international comparison I can find points in the opposite direction for new trial registrations. A 2025 study found that China registered 16,612 trials in 2023, including 7,798 randomized trials, compared with 9,100 trials and 4,619 randomized trials in the United States:

https://www.jclinepi.com/article/S0895-4356%2825%2900124-6/a...

That does not by itself prove that China has a larger stock of trials currently classified as active. It does show that "the US has more active clinical trials than any other country" requires a specific global dataset + status definition + date + deduplication method etc. A count taken only from ClinicalTrials.gov is not a neutral worldwide comparison, because US law requires many FDA regulated trials to be registered there, whereas studies outside its legal and policy scope may be submitted voluntarily:

https://clinicaltrials.gov/policy/fdaaa-801-final-rule

https://clinicaltrials.gov/about

For global comparisons, the WHO's ICTRP is more appropriate because it provides access to ongoing and completed trial records supplied by registries around the world and groups multiple records referring to the same trial:

https://www.who.int/tools/clinical-trials-registry-platform/...

IQVIA's latest result does support a narrower US leadership claim: trial starts became increasingly concentrated among US headquartered sponsors in 2025. But that measures the headquarters of the sponsor (not necessarily the country where the trials took place) and it does not measure the total number of currently active trials:

https://www.iqvia.com/insights/the-iqvia-institute/reports-a...

So the fair conclusion is that the US remains one of the world's largest clinical rial and drug development centres and may lead certain sponsor based or industry sponsored measures. The categorical claim that it has more active clinical trials than every other country has not been demonstrated by your sources.

> In general, I was referring to "major industries" ....

I also do not think the UN classification resolves the "every major industry" question. ISIC is a statistical classification system, not a rule for what ordinary speakers must regard as a single competitive industry.

ISIC Division 30 combines:

shipbuilding railway equipment aircraft and spacecraft military vehicles motorcycles and bicycles

https://unstats.un.org/unsd/publication/seriesm/seriesm_4rev...

Under that aggregation, US aerospace strength can be used to declare the US a "leading player" in the division even though its commercial shipbuilding industry is negligible by global standards. Likewise, ISIC Division 27 combines batteries with motors, generators, wiring, lighting and domestic appliances.

Whenever the US is weak in one globally important industry, it can be bundled with another industry in which the US is strong. Almost any large, diversified economy could be described as a "leading player" in every sufficiently broad category using that method.

> You are correct that the US has "only" 60 GW of domestic solar module production capacity...

The solar comparison also switches metrics. China having more than 80% of global solar module manufacturing capacity is a manufacturing claim. The US being second in solar electricity generation is a deployment/generation claim. Operating solar panels (many of which depend on an overwhelmingly Asian supply chain) does not establish leadership in manufacturing them. The IEA says China has over 80% of module capacity and 95% of wafer capacity.

https://www.iea.org/reports/advancing-clean-technology-manuf...

> True but the US was the 2nd largest producer. In battery cells, we have an estimated capacity of 96 GWh in 2026....

The same issue applies to batteries. The 96 GWh number in your source is projected for 2026, not actual 2025 production. Specifically US energy storage cell capacity, not all lithium ion batteries and capacity, not output.

Your source says US ESS cell capacity was "essentially zero" in 2024 and projects 96 GWh in 2026.

https://poweralliance.org/2026/03/18/american-energy-storage...

Meanwhile, the IEA estimates that China manufactured well over 80% of all batteries in 2025. It says the US and EU each supplied a similar share of the relatively small remainder, while US and European factories remain heavily dependent on imported components. For grid storage LFP batteries specifically, supply is almost entirely Chinese.

https://www.iea.org/commentaries/global-battery-markets-are-...

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#276

If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…

Having tested K3, Qwen 3.8 max preview, Fable and Sol for the past few days, I partially agree. Don’t trust the benchmarks, and the Chinese models really are slow and token-inefficient.

However they do seem very close to SOTA : I’d say roughly equal to previous gen (Opus 4.8, GPT 5.5). It’s yet another silly benchmark, but compare them here: https://senko.net/vibecode-bench/

I also had K3, Qwen3.8 and Fable (using Kimi Code, Qwen Code and Claude Code harnesses respectively, and the official APIs) create a simple but far from trivial web app (zero shot, from a detailed spec). In user testing all three results looked/behaved more or less the same. I had Sol (via codex) do code reviews on all three and it concluded all three were solid, with some room for improvement. Fable was slightly ahead of the pack.

In my own work I still prefer Opus 4.8 (until Antrophic come to their senses and allow 100% Fable usage on Max plans) and Sol, but if I had to find an alternative, I could live with both K3 and Qwen3.8 just fine.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#277

What's the data governance and privacy controls on using Kimi K3 if I subscribe to their coding plans? I want to migrate away from Anthropic

From https://platform.kimi.ai/docs/agreement/modeluse

> We may use Content to provide, maintain, develop, support, and improve the Services, comply with applicable law, enforce our terms and policies, and keep the Services safe and secure. Customer who requires restrictions on the use of Customer Content for training or improving Moonshot AI models may contact Moonshot AI to discuss available enterprise arrangements or separate written agreements. Unless otherwise expressly agreed in writing, Customer Content may be used for the foregoing purposes.

Notably, unlike Claude, there is not an opt-out option for the model training part. The TOS explicitly allows Kimi to train on your code.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#278
post #276

If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…

Having tested K3, Qwen 3.8 max preview, Fable and Sol for the past few days, I partially agree. Don’t trust the benchmarks, and the Chinese models really are slow and token-inefficient. However they do seem very close to SOTA : I’d say roughly equal to previous gen (Opus 4.8, GPT 5.5). It’s yet another silly benchmark, but compare them here: https://senko.net/vibecode-bench/ I also had K3, Qwen3.8 and Fable (using Ki…

Yeah, open weights hasn't fully caught up yet, but it's getting very close. And it's certainly passed the point of being reasonably interchangeable with the frontier for (programming) work. Add in the benefits of not being rug-pulled by the frontier labs silently messing with, the knobs on their models or outright denying you the ability to do certain kinds of work (c.f. the HuggingFace fiasco), and they probably come out ahead in several respects.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#279
post #233

Earlier quoted context omitted.

Yeah. I've found that Opus by default outputs something I call "Claude-lang." It consists of oversimplified, grammatically incomplete sentences that I find painful to read. Maybe it is something that is easy for it to read and write, but definitely not for humans. For example, Skim once now; refer back while reading Part II. \*Every bold technical term in Part II is defined here\* — treat these as a dictionary, not a…

"Grammatically incomplete"?!

It writes in a bunch of annoying too-short sentence fragments, so, yes?

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#280
post #117

Earlier quoted context omitted.

Apparently, the attacker in the Hugging Face case was reported to be an internal OpenAI model trying to break into HF and steal the answers to cybersecurity benchmarks: https://openai.com/index/hugging-face-model-evaluation-secur... It really doesn't matter what restrictions are placed on public use of models if the attacking models are internal models at the AI labs themselves. So if the AI labs are literally runnin…

What? The issue is that models are not well-controlled and are increasingly powerful. Offense/defense/Chinese/American/OAI/HuggingFace – none of it matters. What matters is introducing highly capable intelligences that we - quite demonstrably – do not have effective positive control over.

Oh, to be clear, I don't think that anything about this overall situation is even slightly OK.
Post reply on HN