StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
31–40 of 97 posts
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#32Yet when I tried it it did absymal compared to Gemini 2.5 Flash
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#33Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#34I ran 300+ benchmarks across 15 models in OpenClaw and published two separate leaderboards: performance and cost-effectiveness. The two boards look nothing alike. Top 3 performance: Claude Opus 4.6, GPT-5.4, Claude Sonnet 4.6. Top 3 cost-effectiveness: StepFun 3.5 Flash, Grok 4.1 Fast, MiniMax M2.7. The most dramatic split: Claude Opus 4.6 is #1 on performance but #14 on cost-effectiveness. StepFun 3.5 Flash is #1 co…
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#35Tried the free version on OpenRouter with pi.dev and it's competent at tool calling and creative writing is "good enough" for me (more "natural Claude-level" and not robotic GPT-slop level) but it makes some grave mistakes (had some Hanzi in the output once and typos in words) so it may be good with "simple" agentic workflows but it's definitely not made for programming nor made for long writing.
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#36Pricing is essentially the same: MiMo V2 Flash: $0.09/M input, $0.29/M output Step 3.5 Flash: $0.10/M input, $0.30/M output
MiMo has 41 vs 38 for Step on the Artificial Analysis Intelligence Index, but it's 49 vs 52 for Step on their Agentic Index.
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#37Earlier quoted context omitted.
Could you add a column for time or number of tokens? Some models take forever because of their excessive reasoning chains.
both are shown in battle detail page already. Time is shown in Scores table. Number of tokens are shown in Cost details at the bottom of the Scores. (I thought most people just want to see cost in USD so I put token details at the bottom)
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#38I ran 300+ benchmarks across 15 models in OpenClaw and published two separate leaderboards: performance and cost-effectiveness. The two boards look nothing alike. Top 3 performance: Claude Opus 4.6, GPT-5.4, Claude Sonnet 4.6. Top 3 cost-effectiveness: StepFun 3.5 Flash, Grok 4.1 Fast, MiniMax M2.7. The most dramatic split: Claude Opus 4.6 is #1 on performance but #14 on cost-effectiveness. StepFun 3.5 Flash is #1 co…
Cheapest just isn't a very useful metric. Can I suggest a Pareto-curve type representation? Cost / request vs ELO would be useful and you have all the data.
Essentially I'm using the relative rank in each battle to fit a latent strength for each model, and then use a nonlinear function to map the latent strength to Elo just for human readability. The map function is actually arbitrary as long as it's a monotonically increasing function so it preserves the rank. The only reliable result (that is invariant to the choice of the function) is the relative rank of models.
That being said, if I use score/cost as metrics, the rank completely depends on the function I choose, like I can choose a more super-linear function to make high performance model rank higher in score/cost board, or use a more sub-linear function to make low performance model rank higher.
That's why I eventually tried another (the current) approach: let judge give relative rank of models just by looking at cost-effectiveness (consider both performance and cost), and compute the cost-effectiveness leaderboard directly, so the score mapping function does not affect the leaderboard at all.
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#39Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#40Missing from the comparison is MiMo V2 Flash (not Pro), which I think could put up a good fight against Step 3.5 Flash. Pricing is essentially the same: MiMo V2 Flash: $0.09/M input, $0.29/M output Step 3.5 Flash: $0.10/M input, $0.30/M output MiMo has 41 vs 38 for Step on the Artificial Analysis Intelligence Index, but it's 49 vs 52 for Step on their Agentic Index.