StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
21–30 of 97 posts
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#22According to openrouter.ai it looks like StepFun 3.5 Flash is the most popular model at 3.5T tokens, vs GLM 5 Turbo at 2.5T tokens. Claude Sonnet is in 5th place with 1.05T tokens. Which isn't super suprising as StepFun is ~about 5% the price of Sonnet. https://openrouter.ai/apps?url=https%3A%2F%2Fopenclaw.ai%2F
> the most popular model It was free for a long time. That usually skews the statistics. It was the same with grok-code-fast1.
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#23Earlier quoted context omitted.
all 300+ battle data are available at https://app.uniclaw.ai/arena/battles , every single battle is shown with raw conversional history, produced files, judge's verdict and final scores
Thanks! Is the judge an LLM? There's lot of references to "just like LMArena", but LMArena is human evaluated?
Yes, judge is one of opus 4.6, gpt 5.4, gemini 3.1 pro (submitter can choose). Self judge (judge model is also one of the participants) is excluded when computing ranking.
> There's lot of references to "just like LMArena", but LMArena is human evaluated?
Yeah LMArena is human evaluated, but here i found it not practical to gather enough human evaluation data because the effort it take to compare the result is much higher:
- for code, judge needs to read through it to check code quality, and actually run it to see the output
- when producing a webpage or a document, judge needs to check the content and layout visually
- when anything goes wrong, judge needs to read the execution log to see whether partial credit shall be granted
if you look at the cost details of each battle (available at the bottom of battle detail page), judge typically cost more than any participant model.
if we evaluate with human, i would say each evaluation can easily take ~5-10 min
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#24why do half the comments here read like ai trying to boost some sort of scam?
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#25Earlier quoted context omitted.
Thanks! Is the judge an LLM? There's lot of references to "just like LMArena", but LMArena is human evaluated?
> Is the judge an LLM? Yes, judge is one of opus 4.6, gpt 5.4, gemini 3.1 pro (submitter can choose). Self judge (judge model is also one of the participants) is excluded when computing ranking. > There's lot of references to "just like LMArena", but LMArena is human evaluated? Yeah LMArena is human evaluated, but here i found it not practical to gather enough human evaluation data because the effort it take to compa…
Thanks for replying btw, didn't mean any disrespect, good on you for not getting aggro about feedback
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#26I ran 300+ benchmarks across 15 models in OpenClaw and published two separate leaderboards: performance and cost-effectiveness. The two boards look nothing alike. Top 3 performance: Claude Opus 4.6, GPT-5.4, Claude Sonnet 4.6. Top 3 cost-effectiveness: StepFun 3.5 Flash, Grok 4.1 Fast, MiniMax M2.7. The most dramatic split: Claude Opus 4.6 is #1 on performance but #14 on cost-effectiveness. StepFun 3.5 Flash is #1 co…
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#27Earlier quoted context omitted.
> the most popular model It was free for a long time. That usually skews the statistics. It was the same with grok-code-fast1.
Exactly. When I read the headline I thought: "Ofc it is, its free."
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#28I ran 300+ benchmarks across 15 models in OpenClaw and published two separate leaderboards: performance and cost-effectiveness. The two boards look nothing alike. Top 3 performance: Claude Opus 4.6, GPT-5.4, Claude Sonnet 4.6. Top 3 cost-effectiveness: StepFun 3.5 Flash, Grok 4.1 Fast, MiniMax M2.7. The most dramatic split: Claude Opus 4.6 is #1 on performance but #14 on cost-effectiveness. StepFun 3.5 Flash is #1 co…
This has also been my subjective experience But has also been objective in terms of cost.
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#29I ran 300+ benchmarks across 15 models in OpenClaw and published two separate leaderboards: performance and cost-effectiveness. The two boards look nothing alike. Top 3 performance: Claude Opus 4.6, GPT-5.4, Claude Sonnet 4.6. Top 3 cost-effectiveness: StepFun 3.5 Flash, Grok 4.1 Fast, MiniMax M2.7. The most dramatic split: Claude Opus 4.6 is #1 on performance but #14 on cost-effectiveness. StepFun 3.5 Flash is #1 co…
Could you add a column for time or number of tokens? Some models take forever because of their excessive reasoning chains.
Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
#30Earlier quoted context omitted.
> Is the judge an LLM? Yes, judge is one of opus 4.6, gpt 5.4, gemini 3.1 pro (submitter can choose). Self judge (judge model is also one of the participants) is excluded when computing ranking. > There's lot of references to "just like LMArena", but LMArena is human evaluated? Yeah LMArena is human evaluated, but here i found it not practical to gather enough human evaluation data because the effort it take to compa…
Fair enough, yeah, agent evals are hard especially across N models :/ Thanks for replying btw, didn't mean any disrespect, good on you for not getting aggro about feedback