Earlier quoted context omitted.
I strictly prefer when models ignore any human quirks in my responses. Claude trying to be your friend, saying LOL to your jokes is ridiculous and frankly, harmful
You tell jokes to your model? :D
Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
471–480 of 491 posts
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#472Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#473Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#474Earlier quoted context omitted.
https://xkcd.com/605/
This is such a smooth-brain reply, dude. The capability literally exists, today. You're telling me that it's just SO fantastical that people will be able to catch up? Give me a break. Anthropic and OpenAI are not staffed by demigods.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#475What is routing? How do they decide which is better? The only way to come up with a routing model for your workload is to send queries to all the models and then come up with a way to say which solution was better, often trying out multiple times for the same model + query to account for other statistical errors. This makes you, an ai inference user an unwitting AI company with a non scalable product. The biggest men…
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#476If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source mod…
Yes, they are all benchmaxxed, but the question is how benchmaxxed they are relative to each other. We run an evaluation that only compares models in open-ended multi-agent environments where agents affect each other, primarily testing writing code. It's designed to be less vulnerable because there's no solution set, and it's been pretty effective and tends to rank Chinese models lower than their advertised model car…
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#477Earlier quoted context omitted.
Yes, they are all benchmaxxed, but the question is how benchmaxxed they are relative to each other. We run an evaluation that only compares models in open-ended multi-agent environments where agents affect each other, primarily testing writing code. It's designed to be less vulnerable because there's no solution set, and it's been pretty effective and tends to rank Chinese models lower than their advertised model car…
Although you mention cost, I don't see task cost or total evaluation cost per model in your data. Am I just missing it?
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#478Earlier quoted context omitted.
Although you mention cost, I don't see task cost or total evaluation cost per model in your data. Am I just missing it?
The Efficiency tab at https://gertlabs.com/rankings?mode=oneshot_coding (only have cost data for the coding evaluations)
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#479Earlier quoted context omitted.
Chinese government vs American VCs doesn't equate to "slight different players."
American VCs and American Government mingle at the deepest levels. Yes, slightly different is correct.
Re: Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
#480Earlier quoted context omitted.
Chinese government vs American VCs doesn't equate to "slight different players."
True. The Chinese Communist Party's motto is "Serve the People". [0] They might not always live up to that, but at least the aspiration is there. No such pretense from American VCs. [0] https://en.wikipedia.org/wiki/Serve_the_People