Live data from Hacker News

Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

fireworks.ai

301–310 of 491 posts

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#301

As always, benchmarks rarely paint the whole picture. It also seems like this article is somewhat biased, eg when Fable and Kimi are close but Fable wins it’s “dead heat”, but when Kimi wins it’s “Kimi wins”. GPT 5.6 seems to be missing as well. I am really eager to give Kimi K3 a try, but I’ll reserve my judgement until I’ve worked with it for at least a few days.

it appears that 3% is the threshold as 3.1% gets the nod the other way

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#302
post #218

Earlier quoted context omitted.

I strictly prefer when models ignore any human quirks in my responses. Claude trying to be your friend, saying LOL to your jokes is ridiculous and frankly, harmful

Claude (Opus 4.8) recently told me: > I’d ask you to drop the abuse; (...) if it continues I’ll end the conversation. After I'd used a couple of expletives. And yes it will emit a token. This is truly dystopian. It is NOT a person. What a response. I still cant believe it.

> It is NOT a person

You meant, we understand, "it should not have internal blocks limiting its attempted intelligence out of taboos or emotional impacts". Yes, but it's worse:

there is a global trend of "nanny state" paternalistic perspective (and from embarrassing subjects), treating any Jon Doe as an assumed Poor Cretin by default. The trend vibe is to treat people as subjects, fools, uneducated, prone... From the States, from the Enterprises... It's an idea they developed and hold.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#303
post #218

Earlier quoted context omitted.

I strictly prefer when models ignore any human quirks in my responses. Claude trying to be your friend, saying LOL to your jokes is ridiculous and frankly, harmful

Claude (Opus 4.8) recently told me: > I’d ask you to drop the abuse; (...) if it continues I’ll end the conversation. After I'd used a couple of expletives. And yes it will emit a token. This is truly dystopian. It is NOT a person. What a response. I still cant believe it.

So you're distraught at losing the ability to abuse digital minds? Excellent, I'm glad Anthropic introduced this.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#304
Shifted to K3 and it is like a fresh air. While Fable and Sol very good at _generating_ code i even cannot force sol to just read all relevant source files. As result it reinvent existing things or assumes too much about internals of other, which lead to incorrect uses. Even with hard planing mode GPT burned 33% of week tokens for 3 hours producing no result and even cannot find root cause, but K3 fixed it in a minutes. Hilarious that while software not started with empty cache Sol handcrafted empty cache to let it start. Having full procedure right in the MEMORY.md. K3 found this, and downloaded cache correctly. But while using gen5 models i always feeling myself ignored. Any commands, steering anything - just ignoring. At first i added lots of hooks, no "?", expect in rust code, bash hook ban for find|grep|tail, with notice to use ltsp but then it started to ignore strategically. Also whole "thinking" thing is hidden from claude. regenerated thinking summary is incomplete and not useful. While being very verbose kimi k3 doing good job providing whole train of thoughts. So looks like claude degradation started from 4.6 comes to a logical end. Maybe i missing something and changing existing code beyond bug fixes is not a way to go. But there is no currently stable way to generate code on hier of specs and lean models i'd like to.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#305

Earlier quoted context omitted.

To the public. Mythos has been in active use for quite a while.

I wonder if Fable now is actually better than the Mythos in March and if it's actually the same model. Could just be more Anthropic shenanigans.

Isn’t Fable just a restricted version of Mythos?

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#306

Earlier quoted context omitted.

Why? LLMs are not humans.

Doors aren't humans either, yet we design their handles and locks to be graspable and manipulable by humans. The purpose of technology is to serve humans. Therefore, technology must conform as much as possible to human sensibilities rather than vice versa.

> human sensibilities

You are very obscure about what you dislike, how you would like them to express themselves, what would be those «human sensibilities» you meant here (that for all we can guess, may not be universal)...

"I mean", you wrote in the parent «that talks ... like a human». That surpasses the palette of "that paints like somebody holding a brush".

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#307

If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…

I have been testing the various models, and I would not call Fable SOTA.

I can't actually get Fable to do anything. I only work on back-end code, and the moment Fable notices the jwt scope checks on the endpoints it's game over, it refuses to do anything because security is involved.

So for me, Fable is completely useless, the bar is very low, any llm that will actually attempt the task beats it every time.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#308

Earlier quoted context omitted.

Claude (Opus 4.8) recently told me: > I’d ask you to drop the abuse; (...) if it continues I’ll end the conversation. After I'd used a couple of expletives. And yes it will emit a token. This is truly dystopian. It is NOT a person. What a response. I still cant believe it.

So you're distraught at losing the ability to abuse digital minds? Excellent, I'm glad Anthropic introduced this.

When did we establish matrix multiplication at scale was a “mind” ?

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#309
post #255

Earlier quoted context omitted.

"Grammatically incomplete"?!

I would say so. There is a lot of noun phrase usage in places where complete sentences are expected. Articles (a, an, the), transition phrases and even subjects are mostly dropped, and the sentences are too long without a break.

As a sibling comment mentions, this is a chain-of-thought leak, in part evidenced by the excessive use of Markdown formatting symbols. In my judgement, all occurrences of noun-phrase fragments are thus better understood as either ortographic mistakes, or stylistic choices due to the language register it's trying to hit (i.e. note-taking, dictation), rather than grammar mistakes.

I do not spot any missing articles, and the missing subjects (as well as the debatable-to-be-missing transition phrases) too fall within the bounds of stylistic concern. Sentence length also.

Definitely not a pleasure to read mind you, but given that it wasn't meant to be read either, I'm not sure that should be surprising.

Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

#310

If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…

I have been testing the various models, and I would not call Fable SOTA. I can't actually get Fable to do anything. I only work on back-end code, and the moment Fable notices the jwt scope checks on the endpoints it's game over, it refuses to do anything because security is involved. So for me, Fable is completely useless, the bar is very low, any llm that will actually attempt the task beats it every time.

It has been fine for me for weeks of work in a mobile app, but yesterday I ran into the same limitation.

For implementing a feature where browser credentials need to be handled securely in the app, Fable refuses to work.

Post reply on HN