StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
1–10 of 97 posts
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#2The two boards look nothing alike. Top 3 performance: Claude Opus 4.6, GPT-5.4, Claude Sonnet 4.6. Top 3 cost-effectiveness: StepFun 3.5 Flash, Grok 4.1 Fast, MiniMax M2.7.
The most dramatic split: Claude Opus 4.6 is #1 on performance but #14 on cost-effectiveness. StepFun 3.5 Flash is #1 cost-effectiveness, #5 performance.
Other surprises: GLM-5 Turbo, Xiaomi MiMo v2 Pro, and MiniMax M2.7 all outrank Gemini 3.1 Pro on performance.
Rankings use relative ordering only (not raw scores) fed into a grouped Plackett-Luce model with bootstrap CIs. Same principle as Chatbot Arena — absolute scores are noisy, but "A beat B" is reliable. Full methodology: https://app.uniclaw.ai/arena/leaderboard/methodology?via=hn
I built this as part of OpenClaw Arena — submit any task, pick 2-5 models, a judge agent evaluates in a fresh VM. Public benchmarks are free.
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#3I ran 300+ benchmarks across 15 models in OpenClaw and published two separate leaderboards: performance and cost-effectiveness. The two boards look nothing alike. Top 3 performance: Claude Opus 4.6, GPT-5.4, Claude Sonnet 4.6. Top 3 cost-effectiveness: StepFun 3.5 Flash, Grok 4.1 Fast, MiniMax M2.7. The most dramatic split: Claude Opus 4.6 is #1 on performance but #14 on cost-effectiveness. StepFun 3.5 Flash is #1 co…
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#4I ran 300+ benchmarks across 15 models in OpenClaw and published two separate leaderboards: performance and cost-effectiveness. The two boards look nothing alike. Top 3 performance: Claude Opus 4.6, GPT-5.4, Claude Sonnet 4.6. Top 3 cost-effectiveness: StepFun 3.5 Flash, Grok 4.1 Fast, MiniMax M2.7. The most dramatic split: Claude Opus 4.6 is #1 on performance but #14 on cost-effectiveness. StepFun 3.5 Flash is #1 co…
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#5Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#6I ran 300+ benchmarks across 15 models in OpenClaw and published two separate leaderboards: performance and cost-effectiveness. The two boards look nothing alike. Top 3 performance: Claude Opus 4.6, GPT-5.4, Claude Sonnet 4.6. Top 3 cost-effectiveness: StepFun 3.5 Flash, Grok 4.1 Fast, MiniMax M2.7. The most dramatic split: Claude Opus 4.6 is #1 on performance but #14 on cost-effectiveness. StepFun 3.5 Flash is #1 co…
Please don’t use AI to write comments, it cuts against HN guidelines.
gemini is very unreliable at using skills, often just read skills and decide to do nothing.
stepfun leads cost-effectiveness leaderboard.
ranking really depends on tasks, better try your own task.
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#7Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#8According to openrouter.ai it looks like StepFun 3.5 Flash is the most popular model at 3.5T tokens, vs GLM 5 Turbo at 2.5T tokens. Claude Sonnet is in 5th place with 1.05T tokens. Which isn't super suprising as StepFun is ~about 5% the price of Sonnet. https://openrouter.ai/apps?url=https%3A%2F%2Fopenclaw.ai%2F
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#9According to openrouter.ai it looks like StepFun 3.5 Flash is the most popular model at 3.5T tokens, vs GLM 5 Turbo at 2.5T tokens. Claude Sonnet is in 5th place with 1.05T tokens. Which isn't super suprising as StepFun is ~about 5% the price of Sonnet. https://openrouter.ai/apps?url=https%3A%2F%2Fopenclaw.ai%2F
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#10Earlier quoted context omitted.
Please don’t use AI to write comments, it cuts against HN guidelines.
sorry didn't know that. Here is my hand writing tldr: gemini is very unreliable at using skills, often just read skills and decide to do nothing. stepfun leads cost-effectiveness leaderboard. ranking really depends on tasks, better try your own task.