Tried the free version on OpenRouter with pi.dev and it's competent at tool calling and creative writing is "good enough" for me (more "natural Claude-level" and not robotic GPT-slop level) but it makes some grave mistakes (had some Hanzi in the output once and typos in words) so it may be good with "simple" agentic workflows but it's definitely not made for programming nor made for long writing.
StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
51–60 of 97 posts
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#52I was excited to read through this to find out how these tasks are evaluated at scale. Lots of scary looking formulas with sigmas and other Greek letters. Then I clicked on one task to see what it looks like “on the ground”: https://app.uniclaw.ai/arena/DDquysCGBsHa (not cherry picked- literally the first one I clicked on) The task was: > Find rental properties with 10 bedrooms and 8 or more bathrooms within a 1 hour…
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#53Earlier quoted context omitted.
Since that discussion, they released the base model and a midtrain checkpoint: - https://huggingface.co/stepfun-ai/Step-3.5-Flash-Base - https://huggingface.co/stepfun-ai/Step-3.5-Flash-Base-Midtra... I'm not aware of other AI labs that released base checkpoint for models in this size class. Qwen released some base models for 3.5, but the biggest one is the 35B checkpoint. They also released the entire training pipel…
Tuned Qwen 3.5 27B beats Step 3.5 on almost all benchmarks, so the point about the size class is moot.
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#54Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#55Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#56Earlier quoted context omitted.
both are shown in battle detail page already. Time is shown in Scores table. Number of tokens are shown in Cost details at the bottom of the Scores. (I thought most people just want to see cost in USD so I put token details at the bottom)
some kind of top-level metric like avg tokens/task would be useful. e.g. yes stepfun is 5% the price of sonnet, but does it use 1x, 10x or 1000x more tokens to accomplish similar tasks/median per task. for example I am willing to eat a 20% quality dive from sonnet if the token use is < 10% more than sonnet. if token use is 1000x then that's something I want to know.
also added per battle stats in battle detail page
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#57Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#58Earlier quoted context omitted.
Tuned Qwen 3.5 27B beats Step 3.5 on almost all benchmarks, so the point about the size class is moot.
Benchmarks are not interesting in deciding the "size class". Bigger size means more knowledge. Also, the Qwen 3.5 27B is a dense 27B active parameter model. StepFun 3.5 Flash has 11B active parameters.
Qwen 3.5 27B beats StepFun 3.5 Flash on GPQA Diamond too, so probably no.
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#59This model is free to use, and has been for quite some time on OpenRouter. $0 is pretty hard to beat in terms of cost effectiveness.