If you haven’t heard of it yet there’s some good discussion here: https://news.ycombinator.com/item?id=47069179
StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
11–20 of 97 posts
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#12I ran 300+ benchmarks across 15 models in OpenClaw and published two separate leaderboards: performance and cost-effectiveness. The two boards look nothing alike. Top 3 performance: Claude Opus 4.6, GPT-5.4, Claude Sonnet 4.6. Top 3 cost-effectiveness: StepFun 3.5 Flash, Grok 4.1 Fast, MiniMax M2.7. The most dramatic split: Claude Opus 4.6 is #1 on performance but #14 on cost-effectiveness. StepFun 3.5 Flash is #1 co…
Please don’t use AI to write comments, it cuts against HN guidelines.
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#13Earlier quoted context omitted.
sorry didn't know that. Here is my hand writing tldr: gemini is very unreliable at using skills, often just read skills and decide to do nothing. stepfun leads cost-effectiveness leaderboard. ranking really depends on tasks, better try your own task.
It’s too late once it’s happened. I was curious, then when I saw the site looked vibecoded and you’re commenting with AI, I decided to stop trying to reason through the discrepancies between what was claimed and what’s on the site (ex. 300 battles vs. only a handful in site data).
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#14StepFun is an interesting model. If you haven’t heard of it yet there’s some good discussion here: https://news.ycombinator.com/item?id=47069179
- https://huggingface.co/stepfun-ai/Step-3.5-Flash-Base
- https://huggingface.co/stepfun-ai/Step-3.5-Flash-Base-Midtra...
I'm not aware of other AI labs that released base checkpoint for models in this size class. Qwen released some base models for 3.5, but the biggest one is the 35B checkpoint.
They also released the entire training pipeline:
- https://huggingface.co/datasets/stepfun-ai/Step-3.5-Flash-SF...
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#15Earlier quoted context omitted.
sorry didn't know that. Here is my hand writing tldr: gemini is very unreliable at using skills, often just read skills and decide to do nothing. stepfun leads cost-effectiveness leaderboard. ranking really depends on tasks, better try your own task.
It’s too late once it’s happened. I was curious, then when I saw the site looked vibecoded and you’re commenting with AI, I decided to stop trying to reason through the discrepancies between what was claimed and what’s on the site (ex. 300 battles vs. only a handful in site data).
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#16StepFun is an interesting model. If you haven’t heard of it yet there’s some good discussion here: https://news.ycombinator.com/item?id=47069179
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#17According to openrouter.ai it looks like StepFun 3.5 Flash is the most popular model at 3.5T tokens, vs GLM 5 Turbo at 2.5T tokens. Claude Sonnet is in 5th place with 1.05T tokens. Which isn't super suprising as StepFun is ~about 5% the price of Sonnet. https://openrouter.ai/apps?url=https%3A%2F%2Fopenclaw.ai%2F
It was free for a long time. That usually skews the statistics. It was the same with grok-code-fast1.
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#18Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#19Earlier quoted context omitted.
It’s too late once it’s happened. I was curious, then when I saw the site looked vibecoded and you’re commenting with AI, I decided to stop trying to reason through the discrepancies between what was claimed and what’s on the site (ex. 300 battles vs. only a handful in site data).
Too late for what? For you? maybe. There are many others that are okay with it and it doesn't disminish the quality of the work. Props to the author.
Maybe? :)
> There are many others that are okay with it
Correct.
> and it doesn't disminish the quality of the work.
It does affect incoming people hearing about the work.
I applaud your instinct to defend someone who put in effort. It's one of the most important things we can do.
Another important thing we can do for them is be honest about our own reactions. It's not sunshine and rainbows on its face, but, it is generous. Mostly because A) it takes time B) other people might see red and harangue you for it.
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#20Earlier quoted context omitted.
It’s too late once it’s happened. I was curious, then when I saw the site looked vibecoded and you’re commenting with AI, I decided to stop trying to reason through the discrepancies between what was claimed and what’s on the site (ex. 300 battles vs. only a handful in site data).
all 300+ battle data are available at https://app.uniclaw.ai/arena/battles , every single battle is shown with raw conversional history, produced files, judge's verdict and final scores