As always, benchmarks rarely paint the whole picture. It also seems like this article is somewhat biased, eg when Fable and Kimi are close but Fable wins it’s “dead heat”, but when Kimi wins it’s “Kimi wins”. GPT 5.6 seems to be missing as well. I am really eager to give Kimi K3 a try, but I’ll reserve my judgement until I’ve worked with it for at least a few days.
Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
301–310 of 491 posts
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#302Earlier quoted context omitted.
I strictly prefer when models ignore any human quirks in my responses. Claude trying to be your friend, saying LOL to your jokes is ridiculous and frankly, harmful
Claude (Opus 4.8) recently told me: > I’d ask you to drop the abuse; (...) if it continues I’ll end the conversation. After I'd used a couple of expletives. And yes it will emit a token. This is truly dystopian. It is NOT a person. What a response. I still cant believe it.
You meant, we understand, "it should not have internal blocks limiting its attempted intelligence out of taboos or emotional impacts". Yes, but it's worse:
there is a global trend of "nanny state" paternalistic perspective (and from embarrassing subjects), treating any Jon Doe as an assumed Poor Cretin by default. The trend vibe is to treat people as subjects, fools, uneducated, prone... From the States, from the Enterprises... It's an idea they developed and hold.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#303Earlier quoted context omitted.
I strictly prefer when models ignore any human quirks in my responses. Claude trying to be your friend, saying LOL to your jokes is ridiculous and frankly, harmful
Claude (Opus 4.8) recently told me: > I’d ask you to drop the abuse; (...) if it continues I’ll end the conversation. After I'd used a couple of expletives. And yes it will emit a token. This is truly dystopian. It is NOT a person. What a response. I still cant believe it.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#304Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#305Earlier quoted context omitted.
To the public. Mythos has been in active use for quite a while.
I wonder if Fable now is actually better than the Mythos in March and if it's actually the same model. Could just be more Anthropic shenanigans.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#306Earlier quoted context omitted.
Why? LLMs are not humans.
Doors aren't humans either, yet we design their handles and locks to be graspable and manipulable by humans. The purpose of technology is to serve humans. Therefore, technology must conform as much as possible to human sensibilities rather than vice versa.
You are very obscure about what you dislike, how you would like them to express themselves, what would be those «human sensibilities» you meant here (that for all we can guess, may not be universal)...
"I mean", you wrote in the parent «that talks ... like a human». That surpasses the palette of "that paints like somebody holding a brush".
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#307If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…
I can't actually get Fable to do anything. I only work on back-end code, and the moment Fable notices the jwt scope checks on the endpoints it's game over, it refuses to do anything because security is involved.
So for me, Fable is completely useless, the bar is very low, any llm that will actually attempt the task beats it every time.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#308Earlier quoted context omitted.
Claude (Opus 4.8) recently told me: > I’d ask you to drop the abuse; (...) if it continues I’ll end the conversation. After I'd used a couple of expletives. And yes it will emit a token. This is truly dystopian. It is NOT a person. What a response. I still cant believe it.
So you're distraught at losing the ability to abuse digital minds? Excellent, I'm glad Anthropic introduced this.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#309Earlier quoted context omitted.
"Grammatically incomplete"?!
I would say so. There is a lot of noun phrase usage in places where complete sentences are expected. Articles (a, an, the), transition phrases and even subjects are mostly dropped, and the sentences are too long without a break.
I do not spot any missing articles, and the missing subjects (as well as the debatable-to-be-missing transition phrases) too fall within the bounds of stylistic concern. Sentence length also.
Definitely not a pleasure to read mind you, but given that it wasn't meant to be read either, I'm not sure that should be surprising.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#310If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…
I have been testing the various models, and I would not call Fable SOTA. I can't actually get Fable to do anything. I only work on back-end code, and the moment Fable notices the jwt scope checks on the endpoints it's game over, it refuses to do anything because security is involved. So for me, Fable is completely useless, the bar is very low, any llm that will actually attempt the task beats it every time.
For implementing a feature where browser credentials need to be handled securely in the app, Fable refuses to work.